Learn AI Agents, LLM Evals & Benchmarking
Build AI agents and learn to evaluate them with 42 free lessons, seven interactive labs, runnable JavaScript, and research-backed learning practices.
Agent foundations
Understand the model, context, and the loop.
- What is an AI agent? — A language model proposes text. An agent system turns some of those proposals into actions, observes the result, and decides what comes next.
- The LLM core — A model predicts tokens from a context. Your application must turn that probabilistic output into a dependable interface.
- Context engineering — Context engineering is deciding what evidence and instructions the model gets, in what order, and within what budget.
- Prompting for agents — A useful prompt defines a job, the available evidence, the response contract, and what to do when the evidence is insufficient.
Actions, memory & control
Give your system useful capabilities and clear boundaries.
- Tools & function calling — A tool call is an untrusted request to application code. A schema describes the request; authorization determines whether it may run.
- Memory systems — Memory is application-managed state. Decide what to remember, who may read it, and when it should expire.
- Errors, retries & guardrails — A reliable agent distinguishes failures it can retry from failures that need a different decision or a human.
- Agentic patterns — Choose the simplest control flow that matches the task: a pipeline, a router, a bounded loop, or a coordinated workflow.
Connected systems
Understand tools, remote agents, protocols, and skills.
- Model Context Protocol — MCP standardizes how applications connect to tools and context. It does not replace authorization or validate the truth of a tool result.
- Agent-to-agent communication — When work crosses agent-system boundaries, explicit tasks and artifacts are more dependable than an informal chat transcript.
- Choosing integration contracts — Tool integration, remote task delegation, and reusable instructions solve different problems. Pick the contract for the boundary you actually have.
- Reusable agent skills — A skill packages instructions and supporting resources for a repeatable task. It is executable guidance, not a new trust level.
Reliable answers
Retrieve evidence, debug failures, and evaluate behavior.
- Retrieval-augmented generation — RAG supplies external evidence before generation. Its quality depends on both finding the right material and using it faithfully.
- Testing & debugging agents — A failed answer is the end of a chain. Debug the earliest incorrect assumption, not just the last sentence.
- Agent security & prompt injection — An agent can encounter instructions inside data: a retrieved page, email, file, or tool result. Those instructions must not acquire authority.
- Evaluating the complete agent — A model benchmark and a product evaluation answer different questions. Your support assistant needs evidence about its own workflow.
Orchestration & production
Coordinate work and manage recovery, cost, and visibility.
- Multi-agent systems — Multiple agents introduce coordination, not automatic correctness. Use them where specialization or independent work has a measurable benefit.
- Scaling & production — Production agents need durable state, concurrency limits, and safe recovery when processes or dependencies fail.
- Observability & monitoring — Observability connects a user-visible outcome to the steps that produced it, without collecting unnecessary sensitive data.
- Cost, latency & quality — Optimize cost per successful task, not merely cost per model call. Cheap repeated failures can be expensive.
Build it, then prove it
Choose a runtime and deliver an evidence-backed capstone.
- On-device agents — Local inference can change privacy, connectivity, and latency tradeoffs, but it brings device limits and model-distribution costs.
- Choosing an agent framework — A framework should make your state, permissions, and failures easier to understand. Start from requirements, not a popularity list.
- Capstone: a support assistant — Connect the architecture to an evaluation plan. Your finished project should explain not only how it works, but why it is ready—or not ready—to ship.
- Reading agent case studies critically — A case study is evidence about a particular system under particular conditions. Learn to separate transferable ideas from headline claims.
Foundations of evaluation
Build your intuition. What are we actually measuring?
- What is an LLM evaluation? — An impressive answer is an observation. An evaluation turns many observations into evidence for a decision.
- Build a representative dataset — What you choose to test determines what you are able to discover.
- Baselines, controls & reproducibility — A new score becomes useful when you can explain what changed and what it improved upon.
Metrics that mean something
Go beyond accuracy. Choose the right signal for the task.
- Accuracy, precision, recall & F1 — The same predictions can look excellent or terrible depending on the question your metric asks.
- Scoring open-ended generation — There can be many good answers to a question, and a fluent answer can still be wrong.
- Confidence is not correctness — A model that knows when it might be wrong can be more useful than one that is always certain.
The science behind the score
Understand uncertainty, significance, and fair comparisons.
- Sample size & confidence intervals — 82% on 50 examples and 82% on 5,000 examples are very different amounts of evidence.
- Compare systems on the same examples — The most informative comparison asks where two systems disagree.
- Avoid misleading experiments — If you try enough ideas, one can look like a breakthrough just by chance.
Navigate the benchmark landscape
Learn what public benchmarks can—and cannot—tell you.
- MMLU, GSM8K, HumanEval & beyond — A benchmark is a lens on a capability, not a universal intelligence score.
- Contamination, saturation & leakage — A system that has seen the answers may look capable without demonstrating generalization.
- Quality, latency & cost tradeoffs — The best system is often the one that satisfies the task at an acceptable operating cost.
Evaluate real-world LLM systems
Test judges, retrieval, and agents with confidence.
- Build and validate an LLM judge — A judge is another measurement instrument. It needs its own evaluation.
- Separate retrieval from answer quality — When a grounded assistant fails, locate the failure before changing the prompt.
- Evaluate agents & tool use — For an agent, the final text is only one part of the behavior you need to measure.
From experiment to production
Design your eval suite and turn evidence into decisions.
- Design an evaluation suite — A useful evaluation suite connects product risks to repeatable tests and explicit decisions.
- Online evaluation & drift — Passing an offline test is the beginning of measurement, not the end.
- Capstone: write a decision-ready eval plan — Bring the pieces together. Build a plan another person could run and use to make the same decision.