AgentLearn

Baselines, controls & reproducibility

A new score becomes useful when you can explain what changed and what it improved upon.

Choose an honest baseline

Compare against the current production system and a simple alternative. A keyword lookup might solve a routing problem cheaply. A previous model version provides a deployment baseline. A random baseline helps interpret multiple-choice tasks, but is rarely the only useful comparison.

Hold the environment steady

Use the same dataset, scoring rules, and tool permissions for each system. Record model identifiers, prompt templates, retrieval index versions, generation settings, dates, and execution errors. If the tool budget changes, the experiment measures the whole new configuration, not just the model.

Expect run-to-run variation

Repeated generations can differ, including under nominally deterministic settings in some services. Repeated runs help reveal that variation, but repetitions of one input are not independent samples of new user requests. Keep input-level and generation-level uncertainty distinct.

Worked example

Worked example: Model B resolves 74 of 100 tasks and model A resolves 68. But B receives ten tool calls and A only two. This is evidence about those two system configurations; it does not isolate the effect of the underlying model.

Build a ladder of comparisons

Compare the candidate with at least one simple baseline: a fixed policy lookup, a majority-intent classifier, or the current shipped assistant. A model-only baseline helps identify whether retrieval or tools add value. An ablation removes one component while keeping the rest fixed; it tests a more specific hypothesis than comparing two completely different stacks.

For example, evaluate the same 100 requests with a fixed retrieval pipeline and an agent loop. Record answer quality, unsupported claims, tool calls, latency, and cost. If the loop adds three calls but resolves no additional cases, the extra complexity has not justified itself under that test.

Avoid a weak-baseline victory

A poorly configured baseline makes almost any candidate look impressive. Give each system a reasonable configuration within a disclosed resource budget. Record model versions, prompts, retrieval settings, and whether tuning data was shared fairly. If the candidate gets more attempts or a stronger verifier, report those differences.

Do not discard failed runs only for one system. Define how timeouts, malformed responses, and tool outages count before the comparison.

Your artifact: a comparison table with one baseline, one candidate, the controlled variables, changed variables, and a predeclared decision rule.

Key takeaway

Change one factor when isolating causes; record all factors when comparing systems.

Knowledge check

Which change prevents a clean model-only comparison?

  1. Using the same held-out questions
  2. Giving only the new model access to search tools
  3. Saving both models’ raw outputs
Answer and explanation

Giving only the new model access to search tools

Different tool access introduces another cause of improvement. It can be a valid system comparison if the difference is clearly reported.

Sources

Read this lesson as Markdown

Continue learning