Testing & debugging agents
A failed answer is the end of a chain. Debug the earliest incorrect assumption, not just the last sentence.
Turn incidents into reproducible cases
Capture the input, prompt version, model configuration, selected evidence, tool calls, timings, and final outcome. Redact sensitive information before storage. A useful trace lets you distinguish a bad retrieval result from a malformed tool request or an unsupported generated claim.
How it works
Unit-test deterministic logic such as schema validation and retry budgets. Use integration tests for tool boundaries and persistence. Use evaluation datasets for probabilistic behavior. Replays with recorded tool responses help isolate application changes, but they do not measure a changed external service or model. Keep both fast local checks and appropriately scoped end-to-end evaluations.
A concrete example
The assistant says a return is eligible because the retrieved policy is outdated. Changing answer wording will not fix the root cause. Add a regression case covering the policy's effective date and assert that retrieval filters or ranks the current version correctly.
Apply it to your assistant
Change the second result to true and confirm the failing-case list updates. Before running the exercise, predict the result. Afterward, explain which assumption changed and add one case where the system should refuse, ask for clarification, or escalate.
All exercise inputs and outputs are deterministic teaching examples. No language model is called. Run the same idea against a versioned dataset before making a production claim.
Key takeaway
Trace the decision chain and add a regression case at the failed boundary.
JavaScript exercise: Testing & debugging agents · code experiment
Change the second result to true and confirm the failing-case list updates.
const cases = [
{ id: 'current-policy', passed: true, stage: 'retrieval' },
{ id: 'expired-policy', passed: false, stage: 'retrieval' },
{ id: 'citation-format', passed: true, stage: 'answer' },
];
console.log('Regressions:', cases.filter(test => !test.passed));
console.log({ passed: cases.filter(test => test.passed).length, total: cases.length });
Knowledge check
A correct answer generator receives an outdated policy. Where should the first fix be investigated?
- The retrieval freshness/version boundary
- Only the final answer's tone
- The page animation
Answer and explanation
The retrieval freshness/version boundary
The earliest incorrect input is the outdated evidence. Fixing surface wording leaves the cause in place.
Sources
- OpenTelemetry concepts — OpenTelemetry authors, Living documentation. Traces, metrics, and logs for observing distributed systems.
Continue learning
- Retrieval-augmented generation — RAG supplies external evidence before generation. Its quality depends on both finding the right material and using it faithfully.
- Testing & debugging agents — A failed answer is the end of a chain. Debug the earliest incorrect assumption, not just the last sentence.
- Agent security & prompt injection — An agent can encounter instructions inside data: a retrieved page, email, file, or tool result. Those instructions must not acquire authority.
- Evaluating the complete agent — A model benchmark and a product evaluation answer different questions. Your support assistant needs evidence about its own workflow.