# Avoid misleading experiments

Canonical URL: https://agentlearn.dev/learn/evals/experiment-design
Author: [Hemanth HM](https://h3manth.com)
Track: evals
Reading time: 12 minutes

If you try enough ideas, one can look like a breakthrough just by chance.

## Plan before looking

Write down the primary metric, sample size, slices, baseline, and stopping rule before inspecting test results. If you inspect daily and stop the moment a result looks significant, ordinary fixed-sample inference may no longer have its stated error rate. Use a suitable sequential design if continuous inspection is required.

## Account for many comparisons

Testing dozens of prompts and reporting only the winner creates selection bias. Use development data to choose candidates, then confirm on a holdout. When formally testing many hypotheses, choose an appropriate correction or control procedure and distinguish planned analyses from exploratory discoveries.

## Check what the aggregate hides

A model can improve overall while regressing for a language or workflow. Different mixture weights can even reverse the ranking between datasets. Publish both aggregate results and meaningful slices, with counts and uncertainty. Tiny slices deserve investigation but may not support a strong conclusion.

## Worked example

Worked example: You try 30 prompt variants on a 40-item set. The winner gets 39 correct. That observed maximum reflects both skill and selection. Freeze the winning prompt, then evaluate it on fresh examples before treating 97.5% as a generalization estimate.

## Predeclare what will change

Write a hypothesis before running the comparison: “Adding current policy retrieval reduces unsupported answers without increasing p95 latency beyond our limit.” Identify the unit of analysis, primary metric, meaningful effect size, sample plan, and stopping rule. Keep a record even when the result is disappointing.

Change one factor when you want to attribute causality to that factor. If you change the model, prompt, retrieval, and retry budget together, you are comparing configurations, not isolating why performance changed. Configuration comparisons are useful; label them honestly.

## Handle repeated runs and selection

Randomness in generation introduces within-case variability. Run repeated generations when that variability matters and preserve case-level clustering in the analysis. Record seeds where supported, but do not assume a seed guarantees identical results across environments.

Repeatedly peeking at results and stopping when a threshold is crossed changes error properties. Decide on a fixed sample plan or use a valid sequential method. Testing many prompts and reporting only the winner also creates selection bias; reserve fresh confirmation data.

**Your artifact:** an experiment card with hypothesis, baseline, candidate, versions, sample unit, primary metric, constraints, analysis method, and decision. Add a failure taxonomy before looking at outputs so the categories do not merely explain away the result.


## Key takeaway

Predefine your decision rule and confirm exploratory wins on new data.

## Knowledge check

Why can the best of 30 prompts look artificially good?

1. Selection can favor a prompt that got lucky on the test sample
2. More prompts always make every model worse
3. A high score proves data contamination

Answer: Selection can favor a prompt that got lucky on the test sample

Picking the maximum also selects favorable noise. A fresh holdout helps separate a real improvement from that selection effect.

## Sources

- [Confidence intervals for a binomial proportion](https://www.itl.nist.gov/div898/handbook/prc/section2/prc241.htm) — NIST/SEMATECH, Handbook. Statistical reference for uncertainty in pass/fail measurements, including Wilson intervals.
