# Testing & debugging agents

Canonical URL: https://agentlearn.dev/learn/agents/testing-debugging
Author: [Hemanth HM](https://h3manth.com)
Track: agents
Reading time: 10 minutes

A failed answer is the end of a chain. Debug the earliest incorrect assumption, not just the last sentence.

## Turn incidents into reproducible cases

Capture the input, prompt version, model configuration, selected evidence, tool calls, timings, and final outcome. Redact sensitive information before storage. A useful trace lets you distinguish a bad retrieval result from a malformed tool request or an unsupported generated claim.

## How it works

Unit-test deterministic logic such as schema validation and retry budgets. Use integration tests for tool boundaries and persistence. Use evaluation datasets for probabilistic behavior. Replays with recorded tool responses help isolate application changes, but they do not measure a changed external service or model. Keep both fast local checks and appropriately scoped end-to-end evaluations.

## A concrete example

The assistant says a return is eligible because the retrieved policy is outdated. Changing answer wording will not fix the root cause. Add a regression case covering the policy's effective date and assert that retrieval filters or ranks the current version correctly.

## Apply it to your assistant

Change the second result to true and confirm the failing-case list updates. Before running the exercise, predict the result. Afterward, explain which assumption changed and add one case where the system should refuse, ask for clarification, or escalate.

All exercise inputs and outputs are deterministic teaching examples. No language model is called. Run the same idea against a versioned dataset before making a production claim.


## Key takeaway

Trace the decision chain and add a regression case at the failed boundary.

## JavaScript exercise: Testing & debugging agents · code experiment

Change the second result to true and confirm the failing-case list updates.

```javascript
const cases = [
  { id: 'current-policy', passed: true, stage: 'retrieval' },
  { id: 'expired-policy', passed: false, stage: 'retrieval' },
  { id: 'citation-format', passed: true, stage: 'answer' },
];
console.log('Regressions:', cases.filter(test => !test.passed));
console.log({ passed: cases.filter(test => test.passed).length, total: cases.length });
```

## Knowledge check

A correct answer generator receives an outdated policy. Where should the first fix be investigated?

1. The retrieval freshness/version boundary
2. Only the final answer's tone
3. The page animation

Answer: The retrieval freshness/version boundary

The earliest incorrect input is the outdated evidence. Fixing surface wording leaves the cause in place.

## Sources

- [OpenTelemetry concepts](https://opentelemetry.io/docs/concepts/) — OpenTelemetry authors, Living documentation. Traces, metrics, and logs for observing distributed systems.
