# Errors, retries & guardrails

Canonical URL: https://agentlearn.dev/learn/agents/error-handling
Author: [Hemanth HM](https://h3manth.com)
Track: agents
Reading time: 10 minutes

A reliable agent distinguishes failures it can retry from failures that need a different decision or a human.

## Not every failure means try again

A temporary unavailable response may justify a retry. Invalid arguments require correction. Permission denied requires stopping or an approved alternative. An ambiguous timeout after a write is especially dangerous: the operation may already have succeeded even though the response was lost.

## How it works

Set retry and elapsed-time budgets. Use exponential backoff with jitter to avoid synchronized retries under load. For writes, reuse an idempotency key and query operation status when supported. Record the original error, each attempt, and final outcome. A fallback that invents a successful refund is worse than an honest failure.

## A concrete example

A refund request times out after the payment system accepts it. Retrying with a new operation identity risks paying twice. The exercise simulates a repeated request with the same identity and returns the existing result. It demonstrates the concept, not a concurrency-safe payment implementation.

## Apply it to your assistant

Change the second key to refund-2. Why does the balance change twice? Before running the exercise, predict the result. Afterward, explain which assumption changed and add one case where the system should refuse, ask for clarification, or escalate.

All exercise inputs and outputs are deterministic teaching examples. No language model is called. Run the same idea against a versioned dataset before making a production claim.


## Key takeaway

Retry transient failures within a budget; make repeated writes safe before retrying them.

## JavaScript exercise: Errors, retries & guardrails · code experiment

Change the second key to refund-2. Why does the balance change twice?

```javascript
const results = new Map();
let refunded = 0;
function refund(key, amount) {
  if (results.has(key)) return results.get(key);
  refunded += amount;
  const result = { status: 'accepted', amount };
  results.set(key, result); return result;
}
console.log(refund('refund-1', 20));
console.log(refund('refund-1', 20));
console.log({ refunded });
```

## Knowledge check

A refund request times out after being sent. What is the safest next step?

1. Immediately repeat it with a new key
2. Assume it failed and promise no charge
3. Check status or retry with the same idempotency key

Answer: Check status or retry with the same idempotency key

A timeout does not reveal whether the side effect happened. A stable operation identity lets the receiving service deduplicate a retry.

## Sources

- [OpenTelemetry concepts](https://opentelemetry.io/docs/concepts/) — OpenTelemetry authors, Living documentation. Traces, metrics, and logs for observing distributed systems.
