AgentLearn

Online evaluation & drift

Passing an offline test is the beginning of measurement, not the end.

Monitor the population

Track changes in request topics, languages, document freshness, error types, and tool availability. A stable aggregate score can hide a changing mixture of users. Sample real interactions for review with appropriate data handling and compare meaningful slices over time.

Design online experiments carefully

Randomize at an appropriate unit, such as user or organization, to avoid cross-condition interference. Choose a primary outcome, guardrails, and exposure duration. User clicks or thumbs-up are useful observations but can be noisy proxies for correctness and successful task completion.

Close the learning loop

Investigate failures, update the taxonomy, and create development cases. Reassess the holdout when the target population changes substantially. Use a staged rollout and a clear rollback condition when introducing a new system; compare both operational metrics and quality signals.

Worked example

Worked example: The corpus adds a new policy and the retriever still serves old cached passages. Offline scores on last month’s questions remain high. A document-freshness slice and sampled citation review reveal the regression that a global latency chart misses.

Watch for changing inputs and outcomes

A new product line, policy revision, language mix, or tool dependency can change performance even if the prompt stays fixed. Monitor traffic composition and outcomes by slice, not just the overall average. Changes in abstention or escalation rates may be informative before users report wrong answers.

Latency percentiles reveal tails that an average hides. Track p50 and p95 alongside error rates, token usage, tool failures, and task resolution. Define the measurement window and minimum sample sizes before treating a noisy slice as a trend.

Close the loop safely

When a failure is discovered, preserve a redacted reproducible case, classify the boundary that failed, and add an appropriate regression test. If that case is then used for tuning, it is no longer fresh held-out evidence. Keep a separate confirmation set for the next release.

Use staged rollout and rollback criteria suited to the product's risk. A canary with only a few requests cannot establish safety for rare failure modes. High-consequence actions still need hard controls even when dashboards look healthy.

Your artifact: a monitoring plan with three signals, their denominators, collection and privacy rules, alert thresholds, and the human response each alert triggers.

Key takeaway

Monitor changing inputs and real outcomes, then feed evidence back into development.

Knowledge check

Why might last month’s offline score fail to predict today’s performance?

  1. Offline evaluations can never be useful
  2. The user or document distribution may have changed
  3. A stable average latency proves quality improved
Answer and explanation

The user or document distribution may have changed

Distribution shifts can change task difficulty and validity of evidence. Offline evaluation remains useful when its population and setup match the deployment.

Sources

Read this lesson as Markdown

Continue learning