AgentLearn

Learn AI Agents, LLM Evals & Benchmarking

Build AI agents and learn to evaluate them with 42 free lessons, seven interactive labs, runnable JavaScript, and research-backed learning practices.

Agent foundations

Understand the model, context, and the loop.

  • What is an AI agent? — A language model proposes text. An agent system turns some of those proposals into actions, observes the result, and decides what comes next.
  • The LLM core — A model predicts tokens from a context. Your application must turn that probabilistic output into a dependable interface.
  • Context engineering — Context engineering is deciding what evidence and instructions the model gets, in what order, and within what budget.
  • Prompting for agents — A useful prompt defines a job, the available evidence, the response contract, and what to do when the evidence is insufficient.

Actions, memory & control

Give your system useful capabilities and clear boundaries.

  • Tools & function calling — A tool call is an untrusted request to application code. A schema describes the request; authorization determines whether it may run.
  • Memory systems — Memory is application-managed state. Decide what to remember, who may read it, and when it should expire.
  • Errors, retries & guardrails — A reliable agent distinguishes failures it can retry from failures that need a different decision or a human.
  • Agentic patterns — Choose the simplest control flow that matches the task: a pipeline, a router, a bounded loop, or a coordinated workflow.

Connected systems

Understand tools, remote agents, protocols, and skills.

  • Model Context Protocol — MCP standardizes how applications connect to tools and context. It does not replace authorization or validate the truth of a tool result.
  • Agent-to-agent communication — When work crosses agent-system boundaries, explicit tasks and artifacts are more dependable than an informal chat transcript.
  • Choosing integration contracts — Tool integration, remote task delegation, and reusable instructions solve different problems. Pick the contract for the boundary you actually have.
  • Reusable agent skills — A skill packages instructions and supporting resources for a repeatable task. It is executable guidance, not a new trust level.

Reliable answers

Retrieve evidence, debug failures, and evaluate behavior.

  • Retrieval-augmented generation — RAG supplies external evidence before generation. Its quality depends on both finding the right material and using it faithfully.
  • Testing & debugging agents — A failed answer is the end of a chain. Debug the earliest incorrect assumption, not just the last sentence.
  • Agent security & prompt injection — An agent can encounter instructions inside data: a retrieved page, email, file, or tool result. Those instructions must not acquire authority.
  • Evaluating the complete agent — A model benchmark and a product evaluation answer different questions. Your support assistant needs evidence about its own workflow.

Orchestration & production

Coordinate work and manage recovery, cost, and visibility.

  • Multi-agent systems — Multiple agents introduce coordination, not automatic correctness. Use them where specialization or independent work has a measurable benefit.
  • Scaling & production — Production agents need durable state, concurrency limits, and safe recovery when processes or dependencies fail.
  • Observability & monitoring — Observability connects a user-visible outcome to the steps that produced it, without collecting unnecessary sensitive data.
  • Cost, latency & quality — Optimize cost per successful task, not merely cost per model call. Cheap repeated failures can be expensive.

Build it, then prove it

Choose a runtime and deliver an evidence-backed capstone.

  • On-device agents — Local inference can change privacy, connectivity, and latency tradeoffs, but it brings device limits and model-distribution costs.
  • Choosing an agent framework — A framework should make your state, permissions, and failures easier to understand. Start from requirements, not a popularity list.
  • Capstone: a support assistant — Connect the architecture to an evaluation plan. Your finished project should explain not only how it works, but why it is ready—or not ready—to ship.
  • Reading agent case studies critically — A case study is evidence about a particular system under particular conditions. Learn to separate transferable ideas from headline claims.

Foundations of evaluation

Build your intuition. What are we actually measuring?

Metrics that mean something

Go beyond accuracy. Choose the right signal for the task.

The science behind the score

Understand uncertainty, significance, and fair comparisons.

Navigate the benchmark landscape

Learn what public benchmarks can—and cannot—tell you.

Evaluate real-world LLM systems

Test judges, retrieval, and agents with confidence.

From experiment to production

Design your eval suite and turn evidence into decisions.