# Scaling & production

Canonical URL: https://agentlearn.dev/learn/agents/scaling-production
Author: [Hemanth HM](https://h3manth.com)
Track: agents
Reading time: 10 minutes

Production agents need durable state, concurrency limits, and safe recovery when processes or dependencies fail.

## A process restart should not duplicate a refund

An in-memory loop works until the server restarts mid-run. Persist execution state at meaningful boundaries and give external side effects stable operation identities. Resuming a run should inspect what already happened, not blindly replay every action.

## How it works

Use a queue and concurrency limits to absorb bursts without overwhelming model or tool services. Track deadlines, cancellation, and partial failures. Make checkpoints versioned so software updates can interpret old state. Capacity planning should include worst-case loop length and retries, not just the latency of one successful model call.

## A concrete example

A run records that an order lookup succeeded but a draft response was not produced. On resume, it can reuse the authorized observation if still fresh. A payment operation with unknown status needs reconciliation before another attempt.

## Apply it to your assistant

Change completed to include draft and observe the remaining plan. Before running the exercise, predict the result. Afterward, explain which assumption changed and add one case where the system should refuse, ask for clarification, or escalate.

All exercise inputs and outputs are deterministic teaching examples. No language model is called. Run the same idea against a versioned dataset before making a production claim.


## Key takeaway

Design recovery around persisted state and side-effect identity, then test interrupted runs.

## JavaScript exercise: Scaling & production · code experiment

Change completed to include draft and observe the remaining plan.

```javascript
const checkpoint = { runId: 'run-17', completed: ['lookup'], version: 1 };
const plan = ['lookup', 'draft', 'validate'];
console.log('Resume from:', plan.filter(step => !checkpoint.completed.includes(step)));
console.log('External writes additionally require idempotency and status reconciliation.');
```

## Knowledge check

What is unsafe when resuming a crashed run?

1. Inspecting the last checkpoint
2. Blindly re-executing a write whose outcome is unknown
3. Checking the operation status

Answer: Blindly re-executing a write whose outcome is unknown

The write may have succeeded before the crash. Reconciliation or idempotent retry is needed to avoid duplicate effects.

## Sources

- [LangGraph overview](https://docs.langchain.com/oss/javascript/langgraph/overview) — LangChain, Living documentation. Graph-based orchestration, state, persistence, and long-running workflows.
