Design an evaluation suite
A useful evaluation suite connects product risks to repeatable tests and explicit decisions.
Layer the tests
Use fast deterministic checks for parsing and invariants, curated examples for critical behaviors, and broader sampled datasets for estimated performance. Add adversarial cases for specific failure modes. Keep stress-test results distinct from representative traffic estimates so neither is misinterpreted.
Write release criteria
Specify a minimum primary metric, allowable regression margins, critical failure conditions, and operational limits. Compare candidate and baseline on the same inputs. Decide how inconclusive results are handled before seeing the numbers: gather more data, investigate errors, or retain the existing system.
Maintain the instrument
Version datasets, prompts, scorers, and environments. Review examples when policies change. Add newly observed failures to a regression set without using that same set as fresh evidence of generalization. Periodically validate judge behavior and label quality.
Worked example
Worked example: A support release gate requires no unsupported policy claims on a curated critical set, acceptable grounded-answer performance on a representative holdout, and p95 latency under a documented limit. Passing the curated set alone does not guarantee no future policy errors.
Build layers with different costs
Run deterministic unit checks on every edit: schemas, scoring formulas, route normalization, and permission logic. Run a small regression suite for important known failures. Run broader held-out evaluations for release candidates. Production monitoring then watches behavior under actual traffic, subject to privacy and consent constraints.
These layers answer different questions. Unit tests establish local invariants; an evaluation estimates behavior on a defined sample; monitoring detects changes after deployment. None substitutes for the others.
Make failures actionable
Each case should have a stable ID and a failure category. Store a machine-readable result with system and grader versions, raw outcome, score, and relevant trace references. Produce a human-readable report with counts, uncertainty, changed failures, and the proposed decision.
Avoid a single weighted average that hides unacceptable failures. A release can require a quality floor, no observed critical regressions, and a latency ceiling. Zero observed critical failures is still limited evidence; pair the result with coverage and risk reasoning.
Your artifact: three commands or workflow stages: fast checks, regression evaluation, and release evaluation. Specify which failures block a release and who reviews ambiguous results.
Key takeaway
Tie each evaluation to a behavior, an owner, and a decision rule.
Knowledge check
What should happen to a newly discovered production failure?
- Add a regression case and keep fresh evaluation data separate
- Remove the entire evaluation suite
- Treat that one case as representative of all traffic
Answer and explanation
Add a regression case and keep fresh evaluation data separate
A regression case protects against recurrence. Fresh representative data is still needed for a broad estimate after tuning against known failures.
Sources
- Holistic Evaluation of Language Models — Liang et al., 2022. A framework for evaluating multiple dimensions of model behavior, beyond one headline score.
Continue learning
- Design an evaluation suite — A useful evaluation suite connects product risks to repeatable tests and explicit decisions.
- Online evaluation & drift — Passing an offline test is the beginning of measurement, not the end.
- Capstone: write a decision-ready eval plan — Bring the pieces together. Build a plan another person could run and use to make the same decision.