AgentLearn

Interactive AI Agent & LLM Evaluation Labs

Explore seven hands-on experiments covering confidence intervals, precision and recall, cost tradeoffs, agent traces, context, retrieval, and retries.

Labs use clearly labeled synthetic data. JavaScript is required to adjust inputs and run experiments.

  1. Sample size and confidence intervals
  2. Precision and recall
  3. Quality, latency, and cost
  4. Agent trace
  5. Context budget
  6. Retrieval ranking
  7. Retry and recovery