Contamination, saturation & leakage
A system that has seen the answers may look capable without demonstrating generalization.
Trace routes for leakage
Test content can enter pretraining, fine-tuning, few-shot examples, retrieval indexes, or human prompt development. Exact duplicates are only one route: paraphrases, solutions, and related conversation turns can also leak information. Not every overlap proves memorization, but undisclosed overlap weakens interpretation.
Look beyond a saturated score
When most candidates score near the ceiling, a test may no longer distinguish systems usefully. The remaining errors can be ambiguous or mislabeled. Harder tasks, new held-out cases, and operational measurements can provide a more informative comparison than another decimal place on a saturated metric.
Use evidence proportionately
Dataset publication dates and model training cutoffs are clues, not guarantees of cleanliness. Private or newly constructed tests reduce some risks but still need quality control. Document provenance, deduplicate across splits, and consider fresh variations that preserve the intended capability while changing surface form.
Worked example
Worked example: A retrieval index accidentally contains solved benchmark questions. The assistant copies the answers. Its result measures access to the solutions as much as reasoning ability. Excluding those documents and rerunning changes what the experiment can support.
Leakage has several paths
Training data can contain benchmark questions or close variants. Development prompts can include final-test examples. A retrieval index can accidentally expose reference answers. A human can also tune repeatedly against the final set until it effectively becomes development data. These are different leakage paths and require different controls.
A high score alone does not prove contamination, and a paraphrase does not prove novelty. Document what you can verify and what is unknown. Private or newly collected cases can reduce some risks but may introduce their own sampling and labeling problems.
Protect the evaluation boundary
Keep reference answers out of runtime retrieval unless the task explicitly allows them. Restrict final-set access and log evaluation versions. Group near-duplicates before splitting. If a held-out failure becomes a prompt example, move it into the regression suite and use fresh data for the next independent confirmation.
Imagine 20 of 100 test questions are near-duplicates of development examples. The resulting average may overstate generalization. Reporting only the remaining 80 after seeing results introduces another selection choice; specify exclusion rules independently and report the change transparently.
Your artifact: draw a data-flow map from case collection through prompt tuning, retrieval indexing, evaluation, and reporting. Mark every place where test information could leak into the system being tested.
Key takeaway
Protect the boundary between development information and test evidence.
Knowledge check
Which situation is a direct route for evaluation leakage?
- Reporting an uncertainty interval
- Using a benchmark with difficult questions
- Putting held-out answer keys in the retrieval index
Answer and explanation
Putting held-out answer keys in the retrieval index
The system can retrieve the answers during evaluation, undermining the intended test of independent task performance.
Sources
- Measuring Massive Multitask Language Understanding — Hendrycks et al., 2020. The original MMLU paper: multiple-choice evaluation across 57 subjects.
Continue learning
- MMLU, GSM8K, HumanEval & beyond — A benchmark is a lens on a capability, not a universal intelligence score.
- Contamination, saturation & leakage — A system that has seen the answers may look capable without demonstrating generalization.
- Quality, latency & cost tradeoffs — The best system is often the one that satisfies the task at an acceptable operating cost.