Evaluating the complete agent
A model benchmark and a product evaluation answer different questions. Your support assistant needs evidence about its own workflow.
Start with the release decision
Define what must be true before a change ships: correct policy answers, valid citations, no unauthorized writes, acceptable latency, and an affordable cost per resolved request. Evaluate the complete configuration: prompt, retrieval, model, tools, and stopping rules. Changing any component can change the outcome.
How it works
Build a versioned dataset with ordinary cases, edge cases, unanswerable questions, and adversarial inputs. Keep development examples separate from the final held-out evaluation. Use deterministic graders where possible, calibrated human or model judgments where necessary, and inspect failures by slice. Report denominators and uncertainty rather than only a headline percentage.
A concrete example
A toy run answers 18 of 20 cases correctly. That is 90%, but twenty cases provide limited evidence about rare failures. If two failures involve unauthorized refunds, a high average answer score does not make the release safe. Define critical-failure gates separately.
Apply it to your assistant
Make unauthorizedWrites zero and inspect the release decision. Before running the exercise, predict the result. Afterward, explain which assumption changed and add one case where the system should refuse, ask for clarification, or escalate.
All exercise inputs and outputs are deterministic teaching examples. No language model is called. Run the same idea against a versioned dataset before making a production claim.
Key takeaway
Evaluate the full system against a concrete decision; separate average quality from unacceptable failures.
JavaScript exercise: Evaluating the complete agent · code experiment
Make unauthorizedWrites zero and inspect the release decision.
const report = { correct: 18, total: 20, unauthorizedWrites: 1 };
const release = report.correct / report.total >= 0.85 && report.unauthorizedWrites === 0;
console.log({ accuracy: report.correct / report.total, release });
console.log('Toy thresholds for learning, not universal release requirements.');
Knowledge check
A new prompt raises average correctness but introduces an unauthorized refund. Should a strict no-unauthorized-write gate pass?
- Yes, because the average improved
- No, the critical failure violates an independent gate
- Only if the model benchmark is high
Answer and explanation
No, the critical failure violates an independent gate
Critical safety or business constraints should not be averaged away by improvements on other cases.
Sources
- Holistic Evaluation of Language Models — Liang et al., 2022. A framework for evaluating multiple dimensions of model behavior, beyond one headline score.
Continue learning
- Retrieval-augmented generation — RAG supplies external evidence before generation. Its quality depends on both finding the right material and using it faithfully.
- Testing & debugging agents — A failed answer is the end of a chain. Debug the earliest incorrect assumption, not just the last sentence.
- Agent security & prompt injection — An agent can encounter instructions inside data: a retrieved page, email, file, or tool result. Those instructions must not acquire authority.
- Evaluating the complete agent — A model benchmark and a product evaluation answer different questions. Your support assistant needs evidence about its own workflow.