# AgentLearn > Free learning resources for building AI agents and evaluating LLM systems. 42 lessons and seven interactive labs. Basic JavaScript is the prerequisite. Lessons are available as full HTML without JavaScript and as Markdown. Labs use synthetic teaching data; benchmark descriptions are not live rankings. ## Curriculum - [Learn AI Agents, LLM Evals & Benchmarking](https://agentlearn.dev/): Build AI agents and learn to evaluate them with 42 free lessons, seven interactive labs, runnable JavaScript, and research-backed learning practices. - [Build AI Agents: Free JavaScript Course](https://agentlearn.dev/learn/agents): Learn agent loops, tools, memory, RAG, security, and production practices through 24 lessons with runnable JavaScript and a support-assistant capstone. - [LLM Evaluation & Benchmarking Course](https://agentlearn.dev/learn/evals): Learn datasets, scoring, confidence intervals, benchmark limitations, LLM judges, and release decisions through 18 lessons and interactive experiments. - [Interactive AI Agent & LLM Evaluation Labs](https://agentlearn.dev/labs): Explore seven hands-on experiments covering confidence intervals, precision and recall, cost tradeoffs, agent traces, context, retrieval, and retries. - [LLM Benchmark Explorer: Tasks & Limitations](https://agentlearn.dev/benchmarks): Understand MMLU, GSM8K, HumanEval, SWE-bench, GPQA, and MT-Bench: what each measures, its limitations, and links to official benchmarks and papers. - [AI Agents & Evaluation Research References](https://agentlearn.dev/research): Explore the primary research and official documentation behind AgentLearn's agent lessons, evaluation methods, benchmarks, and learning practices. - [AI Agents & LLM Evaluation Glossary](https://agentlearn.dev/glossary): Understand key AI agent and evaluation terms, from precision, recall, and calibration to RAG, tool contracts, prompt injection, MCP, and agent traces. ## Lessons (Markdown) - [What is an AI agent?](https://agentlearn.dev/learn/agents/what-is-an-agent.md): A language model proposes text. An agent system turns some of those proposals into actions, observes the result, and decides what comes next. - [The LLM core](https://agentlearn.dev/learn/agents/llm-core.md): A model predicts tokens from a context. Your application must turn that probabilistic output into a dependable interface. - [Context engineering](https://agentlearn.dev/learn/agents/context-engineering.md): Context engineering is deciding what evidence and instructions the model gets, in what order, and within what budget. - [Prompting for agents](https://agentlearn.dev/learn/agents/prompting.md): A useful prompt defines a job, the available evidence, the response contract, and what to do when the evidence is insufficient. - [Tools & function calling](https://agentlearn.dev/learn/agents/tools.md): A tool call is an untrusted request to application code. A schema describes the request; authorization determines whether it may run. - [Memory systems](https://agentlearn.dev/learn/agents/memory.md): Memory is application-managed state. Decide what to remember, who may read it, and when it should expire. - [Errors, retries & guardrails](https://agentlearn.dev/learn/agents/error-handling.md): A reliable agent distinguishes failures it can retry from failures that need a different decision or a human. - [Agentic patterns](https://agentlearn.dev/learn/agents/patterns.md): Choose the simplest control flow that matches the task: a pipeline, a router, a bounded loop, or a coordinated workflow. - [Model Context Protocol](https://agentlearn.dev/learn/agents/mcp.md): MCP standardizes how applications connect to tools and context. It does not replace authorization or validate the truth of a tool result. - [Agent-to-agent communication](https://agentlearn.dev/learn/agents/a2a.md): When work crosses agent-system boundaries, explicit tasks and artifacts are more dependable than an informal chat transcript. - [Choosing integration contracts](https://agentlearn.dev/learn/agents/protocols-overview.md): Tool integration, remote task delegation, and reusable instructions solve different problems. Pick the contract for the boundary you actually have. - [Reusable agent skills](https://agentlearn.dev/learn/agents/skills.md): A skill packages instructions and supporting resources for a repeatable task. It is executable guidance, not a new trust level. - [Retrieval-augmented generation](https://agentlearn.dev/learn/agents/rag-retrieval.md): RAG supplies external evidence before generation. Its quality depends on both finding the right material and using it faithfully. - [Testing & debugging agents](https://agentlearn.dev/learn/agents/testing-debugging.md): A failed answer is the end of a chain. Debug the earliest incorrect assumption, not just the last sentence. - [Agent security & prompt injection](https://agentlearn.dev/learn/agents/agent-security.md): An agent can encounter instructions inside data: a retrieved page, email, file, or tool result. Those instructions must not acquire authority. - [Evaluating the complete agent](https://agentlearn.dev/learn/agents/evaluations.md): A model benchmark and a product evaluation answer different questions. Your support assistant needs evidence about its own workflow. - [Multi-agent systems](https://agentlearn.dev/learn/agents/multi-agent.md): Multiple agents introduce coordination, not automatic correctness. Use them where specialization or independent work has a measurable benefit. - [Scaling & production](https://agentlearn.dev/learn/agents/scaling-production.md): Production agents need durable state, concurrency limits, and safe recovery when processes or dependencies fail. - [Observability & monitoring](https://agentlearn.dev/learn/agents/observability.md): Observability connects a user-visible outcome to the steps that produced it, without collecting unnecessary sensitive data. - [Cost, latency & quality](https://agentlearn.dev/learn/agents/cost-optimization.md): Optimize cost per successful task, not merely cost per model call. Cheap repeated failures can be expensive. - [On-device agents](https://agentlearn.dev/learn/agents/on-device-agents.md): Local inference can change privacy, connectivity, and latency tradeoffs, but it brings device limits and model-distribution costs. - [Choosing an agent framework](https://agentlearn.dev/learn/agents/agent-frameworks.md): A framework should make your state, permissions, and failures easier to understand. Start from requirements, not a popularity list. - [Capstone: a support assistant](https://agentlearn.dev/learn/agents/putting-together.md): Connect the architecture to an evaluation plan. Your finished project should explain not only how it works, but why it is ready—or not ready—to ship. - [Reading agent case studies critically](https://agentlearn.dev/learn/agents/case-studies.md): A case study is evidence about a particular system under particular conditions. Learn to separate transferable ideas from headline claims. - [What is an LLM evaluation?](https://agentlearn.dev/learn/evals/what-is-eval.md): An impressive answer is an observation. An evaluation turns many observations into evidence for a decision. - [Build a representative dataset](https://agentlearn.dev/learn/evals/dataset-design.md): What you choose to test determines what you are able to discover. - [Baselines, controls & reproducibility](https://agentlearn.dev/learn/evals/baselines.md): A new score becomes useful when you can explain what changed and what it improved upon. - [Accuracy, precision, recall & F1](https://agentlearn.dev/learn/evals/classification.md): The same predictions can look excellent or terrible depending on the question your metric asks. - [Scoring open-ended generation](https://agentlearn.dev/learn/evals/generation.md): There can be many good answers to a question, and a fluent answer can still be wrong. - [Confidence is not correctness](https://agentlearn.dev/learn/evals/calibration.md): A model that knows when it might be wrong can be more useful than one that is always certain. - [Sample size & confidence intervals](https://agentlearn.dev/learn/evals/uncertainty.md): 82% on 50 examples and 82% on 5,000 examples are very different amounts of evidence. - [Compare systems on the same examples](https://agentlearn.dev/learn/evals/paired-comparison.md): The most informative comparison asks where two systems disagree. - [Avoid misleading experiments](https://agentlearn.dev/learn/evals/experiment-design.md): If you try enough ideas, one can look like a breakthrough just by chance. - [MMLU, GSM8K, HumanEval & beyond](https://agentlearn.dev/learn/evals/benchmark-map.md): A benchmark is a lens on a capability, not a universal intelligence score. - [Contamination, saturation & leakage](https://agentlearn.dev/learn/evals/contamination.md): A system that has seen the answers may look capable without demonstrating generalization. - [Quality, latency & cost tradeoffs](https://agentlearn.dev/learn/evals/cost-frontier.md): The best system is often the one that satisfies the task at an acceptable operating cost. - [Build and validate an LLM judge](https://agentlearn.dev/learn/evals/judges.md): A judge is another measurement instrument. It needs its own evaluation. - [Separate retrieval from answer quality](https://agentlearn.dev/learn/evals/rag.md): When a grounded assistant fails, locate the failure before changing the prompt. - [Evaluate agents & tool use](https://agentlearn.dev/learn/evals/agents.md): For an agent, the final text is only one part of the behavior you need to measure. - [Design an evaluation suite](https://agentlearn.dev/learn/evals/suite.md): A useful evaluation suite connects product risks to repeatable tests and explicit decisions. - [Online evaluation & drift](https://agentlearn.dev/learn/evals/monitoring.md): Passing an offline test is the beginning of measurement, not the end. - [Capstone: write a decision-ready eval plan](https://agentlearn.dev/learn/evals/capstone.md): Bring the pieces together. Build a plan another person could run and use to make the same decision. ## Optional - [Full curriculum](https://agentlearn.dev/llms-full.txt): All lesson text, code exercises, knowledge checks, and source citations in one file. - [Sitemap](https://agentlearn.dev/sitemap.xml): Canonical public HTML URLs. Author: [Hemanth HM](https://h3manth.com). Progress and notes stay in the browser. The optional AI helper sends information only on explicit submission.