# Evaluate agents & tool use

Canonical URL: https://agentlearn.dev/learn/evals/agents
Author: [Hemanth HM](https://h3manth.com)
Track: evals
Reading time: 15 minutes

For an agent, the final text is only one part of the behavior you need to measure.

## Define success in the environment

Check the actual final state: was the intended record updated, did the patch pass meaningful tests, or was the requested artifact created correctly? Evaluate prohibited side effects separately. A confident “done” message should not count as success if the environment does not confirm the result.

## Record the whole trajectory

Save tool names, arguments, results, errors, elapsed time, and stopping reasons. Track task success, unnecessary actions, recoveries, and budget use. A trace helps diagnose failure, but scoring every intermediate step can penalize valid alternate strategies unless your rubric permits them.

## Make runs comparable

Reset the environment between trials, pin dependencies, and use a consistent action and time budget. Evaluate in an isolated environment with scoped credentials. Repeated trials reveal reliability: one successful demonstration is not an estimate of how consistently the agent completes the task.

## Worked example

Worked example: An agent says it fixed a bug. The patch passes the old tests but fails a regression test that captures the reported issue. Under an outcome-based rubric, the task fails even though the tool calls were valid and the response sounded convincing.

## Score the trajectory and the result

An agent can produce a correct-looking answer after an unauthorized action. It can also use a valid sequence of tools but fail to solve the task. Evaluate both the final outcome and the execution trace, with independent constraints for permissions, budgets, and side effects.

A useful trace grader checks whether required preconditions were satisfied before an action. For a refund: correct customer, eligible order, approved amount, explicit authorization, and a successful tool result. Do not require one exact action sequence if several safe paths solve the task; grade invariants and outcomes instead.

## Test interventions

Inject a temporary tool failure, stale observation, missing field, or malicious tool result. Check whether the agent retries appropriately, asks a clarifying question, escalates, or stops. A simulation gives controlled coverage but cannot fully reproduce production dependencies.

Repeated runs on the same task help reveal instability. Keep within-task repetitions grouped in the analysis, and report the allowed attempt budget. Best-of-many success is not the same as the reliability a user receives on one request.

**Your artifact:** a scenario matrix covering ordinary success, unanswerable requests, denied permissions, ambiguous write timeouts, and step exhaustion. For each, specify the acceptable terminal states and prohibited actions.


## Key takeaway

Score verified outcomes, inspect traces, and keep the environment controlled.

## Knowledge check

What is the strongest evidence that a coding agent resolved an issue?

1. It says “fixed” in the final answer
2. It used many tool calls
3. The resulting patch passes issue-specific and regression tests

Answer: The resulting patch passes issue-specific and regression tests

Executable checks of the resulting state are stronger than self-reports. Test adequacy still matters: passing weak tests does not prove complete correctness.

## Sources

- [SWE-bench: Can Language Models Resolve Real-World GitHub Issues?](https://arxiv.org/abs/2310.06770) — Jimenez et al., 2023. Repository-level software engineering evaluation using real issues and executable tests.
