AgentLearn

Contamination, saturation & leakage

A system that has seen the answers may look capable without demonstrating generalization.

Trace routes for leakage

Test content can enter pretraining, fine-tuning, few-shot examples, retrieval indexes, or human prompt development. Exact duplicates are only one route: paraphrases, solutions, and related conversation turns can also leak information. Not every overlap proves memorization, but undisclosed overlap weakens interpretation.

Look beyond a saturated score

When most candidates score near the ceiling, a test may no longer distinguish systems usefully. The remaining errors can be ambiguous or mislabeled. Harder tasks, new held-out cases, and operational measurements can provide a more informative comparison than another decimal place on a saturated metric.

Use evidence proportionately

Dataset publication dates and model training cutoffs are clues, not guarantees of cleanliness. Private or newly constructed tests reduce some risks but still need quality control. Document provenance, deduplicate across splits, and consider fresh variations that preserve the intended capability while changing surface form.

Worked example

Worked example: A retrieval index accidentally contains solved benchmark questions. The assistant copies the answers. Its result measures access to the solutions as much as reasoning ability. Excluding those documents and rerunning changes what the experiment can support.

Leakage has several paths

Training data can contain benchmark questions or close variants. Development prompts can include final-test examples. A retrieval index can accidentally expose reference answers. A human can also tune repeatedly against the final set until it effectively becomes development data. These are different leakage paths and require different controls.

A high score alone does not prove contamination, and a paraphrase does not prove novelty. Document what you can verify and what is unknown. Private or newly collected cases can reduce some risks but may introduce their own sampling and labeling problems.

Protect the evaluation boundary

Keep reference answers out of runtime retrieval unless the task explicitly allows them. Restrict final-set access and log evaluation versions. Group near-duplicates before splitting. If a held-out failure becomes a prompt example, move it into the regression suite and use fresh data for the next independent confirmation.

Imagine 20 of 100 test questions are near-duplicates of development examples. The resulting average may overstate generalization. Reporting only the remaining 80 after seeing results introduces another selection choice; specify exclusion rules independently and report the change transparently.

Your artifact: draw a data-flow map from case collection through prompt tuning, retrieval indexing, evaluation, and reporting. Mark every place where test information could leak into the system being tested.

Key takeaway

Protect the boundary between development information and test evidence.

Knowledge check

Which situation is a direct route for evaluation leakage?

  1. Reporting an uncertainty interval
  2. Using a benchmark with difficult questions
  3. Putting held-out answer keys in the retrieval index
Answer and explanation

Putting held-out answer keys in the retrieval index

The system can retrieve the answers during evaluation, undermining the intended test of independent task performance.

Sources

Read this lesson as Markdown

Continue learning