# Evaluating the complete agent

Canonical URL: https://agentlearn.dev/learn/agents/evaluations
Author: [Hemanth HM](https://h3manth.com)
Track: agents
Reading time: 10 minutes

A model benchmark and a product evaluation answer different questions. Your support assistant needs evidence about its own workflow.

## Start with the release decision

Define what must be true before a change ships: correct policy answers, valid citations, no unauthorized writes, acceptable latency, and an affordable cost per resolved request. Evaluate the complete configuration: prompt, retrieval, model, tools, and stopping rules. Changing any component can change the outcome.

## How it works

Build a versioned dataset with ordinary cases, edge cases, unanswerable questions, and adversarial inputs. Keep development examples separate from the final held-out evaluation. Use deterministic graders where possible, calibrated human or model judgments where necessary, and inspect failures by slice. Report denominators and uncertainty rather than only a headline percentage.

## A concrete example

A toy run answers 18 of 20 cases correctly. That is 90%, but twenty cases provide limited evidence about rare failures. If two failures involve unauthorized refunds, a high average answer score does not make the release safe. Define critical-failure gates separately.

## Apply it to your assistant

Make unauthorizedWrites zero and inspect the release decision. Before running the exercise, predict the result. Afterward, explain which assumption changed and add one case where the system should refuse, ask for clarification, or escalate.

All exercise inputs and outputs are deterministic teaching examples. No language model is called. Run the same idea against a versioned dataset before making a production claim.


## Key takeaway

Evaluate the full system against a concrete decision; separate average quality from unacceptable failures.

## JavaScript exercise: Evaluating the complete agent · code experiment

Make unauthorizedWrites zero and inspect the release decision.

```javascript
const report = { correct: 18, total: 20, unauthorizedWrites: 1 };
const release = report.correct / report.total >= 0.85 && report.unauthorizedWrites === 0;
console.log({ accuracy: report.correct / report.total, release });
console.log('Toy thresholds for learning, not universal release requirements.');
```

## Knowledge check

A new prompt raises average correctness but introduces an unauthorized refund. Should a strict no-unauthorized-write gate pass?

1. Yes, because the average improved
2. No, the critical failure violates an independent gate
3. Only if the model benchmark is high

Answer: No, the critical failure violates an independent gate

Critical safety or business constraints should not be averaged away by improvements on other cases.

## Sources

- [Holistic Evaluation of Language Models](https://arxiv.org/abs/2211.09110) — Liang et al., 2022. A framework for evaluating multiple dimensions of model behavior, beyond one headline score.
