# Compare systems on the same examples

Canonical URL: https://agentlearn.dev/learn/evals/paired-comparison
Author: [Hemanth HM](https://h3manth.com)
Track: evals
Reading time: 13 minutes

The most informative comparison asks where two systems disagree.

## Pair by input

Run A and B on the same held-out items. For each item, retain both scores and calculate a difference. Shared easy and hard examples affect both systems, so paired analysis can remove variation caused by item difficulty. A comparison of two unrelated averages throws this information away.

## Estimate the difference directly

For a paired bootstrap, resample input indices with replacement and keep each A/B pair together. Recompute the mean difference in each resample to form an uncertainty distribution. When examples come in correlated groups, resample the groups instead. For binary paired outcomes, McNemar’s test examines the discordant counts.

## Decide what improvement matters

A small p-value does not measure effect size, business value, or the probability the hypothesis is true. Define a practically meaningful improvement in advance, and report the estimated difference with an interval. Overlapping individual model confidence intervals are not a valid replacement for a paired comparison.

## Worked example

Worked example: On 100 fictional tasks, both systems pass 65, only A passes 9, only B passes 15, and neither passes 11. B improves by (15−9)/100 = 6 percentage points. The uncertainty depends on the disagreement pattern, not only the two overall pass rates.

## Preserve which cases each system solved

Suppose two systems run on the same 100 cases. Both pass 70; only A passes 10; only B passes 15; neither passes 5. A scores 80%, B scores 85%, and the observed paired improvement is five percentage points.

The disagreement pattern contains information that two separate totals hide. If cases are matched, do not analyze the difference as though the systems were evaluated on unrelated samples. For binary outcomes, a paired test such as McNemar's focuses on discordant cases. A confidence interval for the paired difference is often more useful than a significance label alone.

## Bootstrap pairs, not individual system scores

Create a row for each case containing both A and B outcomes. Sample rows with replacement, preserving the pair, and recompute the average difference for each resample. Use the distribution to estimate uncertainty with an appropriate interval method. If examples are clustered by conversation, resample conversations rather than individual rows.

If you bootstrap A and B independently, you destroy the pairing. If you choose the best of many candidates on the same test set, a nominal interval does not account for that search.

**Your artifact:** report the paired difference, its uncertainty, the disagreement counts, and the concrete failures that changed. A statistically detectable improvement can still be too small to justify added cost.


## Key takeaway

Analyze paired differences and distinguish statistical evidence from practical value.

## Knowledge check

In a paired bootstrap, what should be resampled together?

1. All A scores separately from all B scores
2. Only the examples where B wins
3. Both systems’ scores for each sampled input

Answer: Both systems’ scores for each sampled input

Keeping the scores paired preserves the relationship between systems on each input. Independent resampling loses that structure.

## Sources

- [Bootstrap Methods: Another Look at the Jackknife](https://doi.org/10.1214/aos/1176344552) — Bradley Efron, 1979. The foundational paper introducing bootstrap resampling for estimating sampling distributions.
