# Observability & monitoring

Canonical URL: https://agentlearn.dev/learn/agents/observability
Author: [Hemanth HM](https://h3manth.com)
Track: agents
Reading time: 10 minutes

Observability connects a user-visible outcome to the steps that produced it, without collecting unnecessary sensitive data.

## Measure outcomes and the path to them

A request-level trace links retrieval, generation, tool execution, and validation. Metrics summarize distributions such as latency, error rate, and cost. Logs capture selected events. These signals complement each other: an average latency number cannot explain why a specific request stalled.

## How it works

Use stable run and span IDs, record model and prompt versions, and label errors consistently. Redact or avoid personal data in prompts and tool results. Record enough metadata to compare releases by task slice. Monitor completion, escalation, critical failures, and user resolution separately; returning HTTP 200 is not the same as solving the request.

## A concrete example

A toy trace spends 80 ms retrieving, 900 ms generating, and 20 ms validating. Optimizing retrieval by 50% saves only 40 ms of the 1,000 ms total. The trace tells you which work dominates, while quality checks tell you whether a faster configuration is acceptable.

## Apply it to your assistant

Double generation time and inspect the new share of total latency. Before running the exercise, predict the result. Afterward, explain which assumption changed and add one case where the system should refuse, ask for clarification, or escalate.

All exercise inputs and outputs are deterministic teaching examples. No language model is called. Run the same idea against a versioned dataset before making a production claim.


## Key takeaway

Link traces to product outcomes, and treat telemetry as sensitive data with a retention policy.

## JavaScript exercise: Observability & monitoring · code experiment

Double generation time and inspect the new share of total latency.

```javascript
const spans = [{ name: 'retrieve', ms: 80 }, { name: 'generate', ms: 900 }, { name: 'validate', ms: 20 }];
const total = spans.reduce((sum, span) => sum + span.ms, 0);
for (const span of spans) console.log({ ...span, share: (100 * span.ms / total).toFixed(1) + '%' });
```

## Knowledge check

Every request returns HTTP 200, but many answers are wrong. What is missing?

1. Task-level quality and outcome measurements
2. More successful HTTP responses
3. A lower logging threshold alone

Answer: Task-level quality and outcome measurements

Transport success does not measure answer correctness or task resolution. Product-level outcome checks are needed.

## Sources

- [OpenTelemetry concepts](https://opentelemetry.io/docs/concepts/) — OpenTelemetry authors, Living documentation. Traces, metrics, and logs for observing distributed systems.
