Observability & monitoring
Observability connects a user-visible outcome to the steps that produced it, without collecting unnecessary sensitive data.
Measure outcomes and the path to them
A request-level trace links retrieval, generation, tool execution, and validation. Metrics summarize distributions such as latency, error rate, and cost. Logs capture selected events. These signals complement each other: an average latency number cannot explain why a specific request stalled.
How it works
Use stable run and span IDs, record model and prompt versions, and label errors consistently. Redact or avoid personal data in prompts and tool results. Record enough metadata to compare releases by task slice. Monitor completion, escalation, critical failures, and user resolution separately; returning HTTP 200 is not the same as solving the request.
A concrete example
A toy trace spends 80 ms retrieving, 900 ms generating, and 20 ms validating. Optimizing retrieval by 50% saves only 40 ms of the 1,000 ms total. The trace tells you which work dominates, while quality checks tell you whether a faster configuration is acceptable.
Apply it to your assistant
Double generation time and inspect the new share of total latency. Before running the exercise, predict the result. Afterward, explain which assumption changed and add one case where the system should refuse, ask for clarification, or escalate.
All exercise inputs and outputs are deterministic teaching examples. No language model is called. Run the same idea against a versioned dataset before making a production claim.
Key takeaway
Link traces to product outcomes, and treat telemetry as sensitive data with a retention policy.
JavaScript exercise: Observability & monitoring · code experiment
Double generation time and inspect the new share of total latency.
const spans = [{ name: 'retrieve', ms: 80 }, { name: 'generate', ms: 900 }, { name: 'validate', ms: 20 }];
const total = spans.reduce((sum, span) => sum + span.ms, 0);
for (const span of spans) console.log({ ...span, share: (100 * span.ms / total).toFixed(1) + '%' });
Knowledge check
Every request returns HTTP 200, but many answers are wrong. What is missing?
- Task-level quality and outcome measurements
- More successful HTTP responses
- A lower logging threshold alone
Answer and explanation
Task-level quality and outcome measurements
Transport success does not measure answer correctness or task resolution. Product-level outcome checks are needed.
Sources
- OpenTelemetry concepts — OpenTelemetry authors, Living documentation. Traces, metrics, and logs for observing distributed systems.
Continue learning
- Multi-agent systems — Multiple agents introduce coordination, not automatic correctness. Use them where specialization or independent work has a measurable benefit.
- Scaling & production — Production agents need durable state, concurrency limits, and safe recovery when processes or dependencies fail.
- Observability & monitoring — Observability connects a user-visible outcome to the steps that produced it, without collecting unnecessary sensitive data.
- Cost, latency & quality — Optimize cost per successful task, not merely cost per model call. Cheap repeated failures can be expensive.