Baselines, controls & reproducibility
A new score becomes useful when you can explain what changed and what it improved upon.
Choose an honest baseline
Compare against the current production system and a simple alternative. A keyword lookup might solve a routing problem cheaply. A previous model version provides a deployment baseline. A random baseline helps interpret multiple-choice tasks, but is rarely the only useful comparison.
Hold the environment steady
Use the same dataset, scoring rules, and tool permissions for each system. Record model identifiers, prompt templates, retrieval index versions, generation settings, dates, and execution errors. If the tool budget changes, the experiment measures the whole new configuration, not just the model.
Expect run-to-run variation
Repeated generations can differ, including under nominally deterministic settings in some services. Repeated runs help reveal that variation, but repetitions of one input are not independent samples of new user requests. Keep input-level and generation-level uncertainty distinct.
Worked example
Worked example: Model B resolves 74 of 100 tasks and model A resolves 68. But B receives ten tool calls and A only two. This is evidence about those two system configurations; it does not isolate the effect of the underlying model.
Build a ladder of comparisons
Compare the candidate with at least one simple baseline: a fixed policy lookup, a majority-intent classifier, or the current shipped assistant. A model-only baseline helps identify whether retrieval or tools add value. An ablation removes one component while keeping the rest fixed; it tests a more specific hypothesis than comparing two completely different stacks.
For example, evaluate the same 100 requests with a fixed retrieval pipeline and an agent loop. Record answer quality, unsupported claims, tool calls, latency, and cost. If the loop adds three calls but resolves no additional cases, the extra complexity has not justified itself under that test.
Avoid a weak-baseline victory
A poorly configured baseline makes almost any candidate look impressive. Give each system a reasonable configuration within a disclosed resource budget. Record model versions, prompts, retrieval settings, and whether tuning data was shared fairly. If the candidate gets more attempts or a stronger verifier, report those differences.
Do not discard failed runs only for one system. Define how timeouts, malformed responses, and tool outages count before the comparison.
Your artifact: a comparison table with one baseline, one candidate, the controlled variables, changed variables, and a predeclared decision rule.
Key takeaway
Change one factor when isolating causes; record all factors when comparing systems.
Knowledge check
Which change prevents a clean model-only comparison?
- Using the same held-out questions
- Giving only the new model access to search tools
- Saving both models’ raw outputs
Answer and explanation
Giving only the new model access to search tools
Different tool access introduces another cause of improvement. It can be a valid system comparison if the difference is clearly reported.
Sources
- Holistic Evaluation of Language Models — Liang et al., 2022. A framework for evaluating multiple dimensions of model behavior, beyond one headline score.
Continue learning
- What is an LLM evaluation? — An impressive answer is an observation. An evaluation turns many observations into evidence for a decision.
- Build a representative dataset — What you choose to test determines what you are able to discover.
- Baselines, controls & reproducibility — A new score becomes useful when you can explain what changed and what it improved upon.