Capstone: write a decision-ready eval plan
Bring the pieces together. Build a plan another person could run and use to make the same decision.
Specify the experiment
Choose an application and write the deployment decision in one sentence. Describe the current baseline and candidate system. List the target population, sampling unit, dataset sources, holdout boundary, and important slices. Define your primary outcome, rubric, uncertainty method, and minimum useful effect.
Plan diagnosis and operation
Include a failure taxonomy and a small set of diagnostic interventions. Explain how you will validate labels or judges. Set time and cost budgets, pin the system configuration, and retain raw outputs. Write a release gate that distinguishes success, failure, and inconclusive evidence.
Communicate the result
Your final report should include the decision, paired score difference and interval, sample counts, slice regressions, costs, and known limitations. Distinguish measured results from assumptions. Save your plan in the notebook and export it as the starting point for a real evaluation project.
Worked example
Capstone prompt: Your team wants to replace a support assistant with a cheaper model. Draft a non-inferiority plan: define the largest acceptable quality loss, protect critical policy cases, measure cost per completed task, and describe what evidence would justify rollout. Do not choose the margin after viewing results.
Produce a reproducible release packet
Use the support assistant from the agents track. Keep the first version read-only. Create a corpus with ordinary return rules and exceptions, then a dataset containing answerable, ambiguous, unanswerable, and adversarial questions. Give each case expected behavior and evidence references.
Version the entire configuration: corpus, chunking, retrieval, prompt, model, decoding, tools, and grader. Compare a fixed workflow baseline with your candidate on matched cases. Record correctness, support, appropriate abstention, critical failures, latency, and cost.
Make the decision, including uncertainty
Your report should contain:
- The product claim and intended population.
- Dataset provenance, slices, split rules, and known coverage gaps.
- Baseline and candidate configurations with identical resource rules where appropriate.
- Per-case results, aggregate metrics, paired differences, and uncertainty.
- The largest regressions and critical failure analysis.
- A release, limited rollout, or no-release decision with monitoring and rollback criteria.
Do not select thresholds after seeing the final results and then present them as predeclared. If evidence is insufficient, “collect more representative cases” is a valid outcome.
Final challenge: ask another developer to reproduce your report using only the packet. If they need your memory to identify the prompt version, scoring rule, or exclusions, improve the packet before declaring the evaluation complete.
Key takeaway
A strong evaluation ends with a defensible decision and a reproducible record.
Knowledge check
When should you choose the largest acceptable quality regression?
- After seeing which margin lets the new model pass
- Only after deployment
- Before inspecting the comparison results
Answer and explanation
Before inspecting the comparison results
Predefining a meaningful margin keeps the release criterion tied to product needs rather than adapting it to a desired result.
Sources
- Holistic Evaluation of Language Models — Liang et al., 2022. A framework for evaluating multiple dimensions of model behavior, beyond one headline score.
Continue learning
- Design an evaluation suite — A useful evaluation suite connects product risks to repeatable tests and explicit decisions.
- Online evaluation & drift — Passing an offline test is the beginning of measurement, not the end.
- Capstone: write a decision-ready eval plan — Bring the pieces together. Build a plan another person could run and use to make the same decision.