AI Agents & Evaluation Research References
Explore the primary research and official documentation behind AgentLearn's agent lessons, evaluation methods, benchmarks, and learning practices.
- ReAct: Synergizing Reasoning and Acting in Language Models — Yao et al., 2022. A research starting point for interleaving model reasoning and actions.
- Attention Is All You Need — Vaswani et al., 2017. The Transformer architecture underlying many language models.
- Lost in the Middle: How Language Models Use Long Contexts — Liu et al., 2023. Evidence that more context does not guarantee effective use of relevant information.
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — Lewis et al., 2020. A foundational approach to combining retrieval with text generation.
- Model Context Protocol specification — MCP contributors, 2026-07-28. Versioned protocol reference. Implementation details should be checked against the version you deploy.
- Agent2Agent Protocol specification — A2A contributors, Living specification. Discovery, messages, tasks, artifacts, and interoperability between agent systems.
- Agent Skills specification — Agent Skills contributors, Living specification. The file format for discoverable instructions and supporting resources.
- WebLLM documentation — MLC AI, Living documentation. Practical documentation for running supported language models in the browser.
- LangGraph overview — LangChain, Living documentation. Graph-based orchestration, state, persistence, and long-running workflows.
- OpenTelemetry concepts — OpenTelemetry authors, Living documentation. Traces, metrics, and logs for observing distributed systems.
- Not what you’ve signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection — Greshake et al., 2023. Research on attacks delivered through content processed by LLM applications.
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering — Yang et al., 2024. A concrete research case study in agent interfaces and executable software tasks.
- Holistic Evaluation of Language Models — Liang et al., 2022. A framework for evaluating multiple dimensions of model behavior, beyond one headline score.
- Measuring Massive Multitask Language Understanding — Hendrycks et al., 2020. The original MMLU paper: multiple-choice evaluation across 57 subjects.
- Evaluating Large Language Models Trained on Code — Chen et al., 2021. HumanEval, functional correctness, and the pass@k estimator.
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena — Zheng et al., 2023. Model-based judging, human agreement, and position and verbosity biases.
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues? — Jimenez et al., 2023. Repository-level software engineering evaluation using real issues and executable tests.
- Training Verifiers to Solve Math Word Problems — Cobbe et al., 2021. GSM8K and the use of verifiers for multi-step mathematical problem solving.
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark — Rein et al., 2023. Expert-written questions in biology, physics, and chemistry.
- Confidence intervals for a binomial proportion — NIST/SEMATECH, Handbook. Statistical reference for uncertainty in pass/fail measurements, including Wilson intervals.
- Bootstrap Methods: Another Look at the Jackknife — Bradley Efron, 1979. The foundational paper introducing bootstrap resampling for estimating sampling distributions.
- More accurate tests for the statistical significance of result differences — Alexander Yeh, 2000. Why dependence between evaluation results matters when testing differences in NLP metrics.
- The science of effective learning with spacing and retrieval practice — Carpenter, Pan & Butler, 2022. The learning research behind recall questions, explanatory feedback, and returning to a topic over time.