# Design an evaluation suite

Canonical URL: https://agentlearn.dev/learn/evals/suite
Author: [Hemanth HM](https://h3manth.com)
Track: evals
Reading time: 13 minutes

A useful evaluation suite connects product risks to repeatable tests and explicit decisions.

## Layer the tests

Use fast deterministic checks for parsing and invariants, curated examples for critical behaviors, and broader sampled datasets for estimated performance. Add adversarial cases for specific failure modes. Keep stress-test results distinct from representative traffic estimates so neither is misinterpreted.

## Write release criteria

Specify a minimum primary metric, allowable regression margins, critical failure conditions, and operational limits. Compare candidate and baseline on the same inputs. Decide how inconclusive results are handled before seeing the numbers: gather more data, investigate errors, or retain the existing system.

## Maintain the instrument

Version datasets, prompts, scorers, and environments. Review examples when policies change. Add newly observed failures to a regression set without using that same set as fresh evidence of generalization. Periodically validate judge behavior and label quality.

## Worked example

Worked example: A support release gate requires no unsupported policy claims on a curated critical set, acceptable grounded-answer performance on a representative holdout, and p95 latency under a documented limit. Passing the curated set alone does not guarantee no future policy errors.

## Build layers with different costs

Run deterministic unit checks on every edit: schemas, scoring formulas, route normalization, and permission logic. Run a small regression suite for important known failures. Run broader held-out evaluations for release candidates. Production monitoring then watches behavior under actual traffic, subject to privacy and consent constraints.

These layers answer different questions. Unit tests establish local invariants; an evaluation estimates behavior on a defined sample; monitoring detects changes after deployment. None substitutes for the others.

## Make failures actionable

Each case should have a stable ID and a failure category. Store a machine-readable result with system and grader versions, raw outcome, score, and relevant trace references. Produce a human-readable report with counts, uncertainty, changed failures, and the proposed decision.

Avoid a single weighted average that hides unacceptable failures. A release can require a quality floor, no observed critical regressions, and a latency ceiling. Zero observed critical failures is still limited evidence; pair the result with coverage and risk reasoning.

**Your artifact:** three commands or workflow stages: fast checks, regression evaluation, and release evaluation. Specify which failures block a release and who reviews ambiguous results.


## Key takeaway

Tie each evaluation to a behavior, an owner, and a decision rule.

## Knowledge check

What should happen to a newly discovered production failure?

1. Add a regression case and keep fresh evaluation data separate
2. Remove the entire evaluation suite
3. Treat that one case as representative of all traffic

Answer: Add a regression case and keep fresh evaluation data separate

A regression case protects against recurrence. Fresh representative data is still needed for a broad estimate after tuning against known failures.

## Sources

- [Holistic Evaluation of Language Models](https://arxiv.org/abs/2211.09110) — Liang et al., 2022. A framework for evaluating multiple dimensions of model behavior, beyond one headline score.
