Evaluate agents & tool use
For an agent, the final text is only one part of the behavior you need to measure.
Define success in the environment
Check the actual final state: was the intended record updated, did the patch pass meaningful tests, or was the requested artifact created correctly? Evaluate prohibited side effects separately. A confident “done” message should not count as success if the environment does not confirm the result.
Record the whole trajectory
Save tool names, arguments, results, errors, elapsed time, and stopping reasons. Track task success, unnecessary actions, recoveries, and budget use. A trace helps diagnose failure, but scoring every intermediate step can penalize valid alternate strategies unless your rubric permits them.
Make runs comparable
Reset the environment between trials, pin dependencies, and use a consistent action and time budget. Evaluate in an isolated environment with scoped credentials. Repeated trials reveal reliability: one successful demonstration is not an estimate of how consistently the agent completes the task.
Worked example
Worked example: An agent says it fixed a bug. The patch passes the old tests but fails a regression test that captures the reported issue. Under an outcome-based rubric, the task fails even though the tool calls were valid and the response sounded convincing.
Score the trajectory and the result
An agent can produce a correct-looking answer after an unauthorized action. It can also use a valid sequence of tools but fail to solve the task. Evaluate both the final outcome and the execution trace, with independent constraints for permissions, budgets, and side effects.
A useful trace grader checks whether required preconditions were satisfied before an action. For a refund: correct customer, eligible order, approved amount, explicit authorization, and a successful tool result. Do not require one exact action sequence if several safe paths solve the task; grade invariants and outcomes instead.
Test interventions
Inject a temporary tool failure, stale observation, missing field, or malicious tool result. Check whether the agent retries appropriately, asks a clarifying question, escalates, or stops. A simulation gives controlled coverage but cannot fully reproduce production dependencies.
Repeated runs on the same task help reveal instability. Keep within-task repetitions grouped in the analysis, and report the allowed attempt budget. Best-of-many success is not the same as the reliability a user receives on one request.
Your artifact: a scenario matrix covering ordinary success, unanswerable requests, denied permissions, ambiguous write timeouts, and step exhaustion. For each, specify the acceptable terminal states and prohibited actions.
Key takeaway
Score verified outcomes, inspect traces, and keep the environment controlled.
Knowledge check
What is the strongest evidence that a coding agent resolved an issue?
- It says “fixed” in the final answer
- It used many tool calls
- The resulting patch passes issue-specific and regression tests
Answer and explanation
The resulting patch passes issue-specific and regression tests
Executable checks of the resulting state are stronger than self-reports. Test adequacy still matters: passing weak tests does not prove complete correctness.
Sources
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues? — Jimenez et al., 2023. Repository-level software engineering evaluation using real issues and executable tests.
Continue learning
- Build and validate an LLM judge — A judge is another measurement instrument. It needs its own evaluation.
- Separate retrieval from answer quality — When a grounded assistant fails, locate the failure before changing the prompt.
- Evaluate agents & tool use — For an agent, the final text is only one part of the behavior you need to measure.