AI Agents & LLM Evaluation Glossary
Understand key AI agent and evaluation terms, from precision, recall, and calibration to RAG, tool contracts, prompt injection, MCP, and agent traces.
- Ablation
- An experiment that removes or changes one component to understand its contribution.
- Accuracy
- The fraction of predictions that are correct: (TP + TN) / all examples.
- Benchmark
- A standardized dataset and evaluation protocol for comparing systems on a defined task.
- Calibration
- Agreement between predicted probabilities and observed frequencies of correctness.
- Confidence interval
- An interval from a procedure designed to cover a fixed population parameter at a stated long-run rate.
- Contamination
- Exposure to test questions, answers, or close equivalents during training or development.
- Coverage
- The fraction of requests a system chooses to answer rather than abstain or defer.
- Distribution shift
- A change in the input or outcome distribution between evaluation and deployment.
- F1 score
- The harmonic mean of precision and recall: 2PR / (P + R). It does not account for true negatives.
- Groundedness
- The extent to which an answer’s claims are supported by the provided evidence.
- Holdout
- Data kept separate from development choices for a less biased evaluation of the finished system.
- LLM-as-a-judge
- A language model used to apply a rubric or compare outputs; the judge needs independent validation.
- Non-inferiority
- Testing whether a candidate is no worse than a baseline by more than a predefined meaningful margin.
- Paired bootstrap
- Resampling inputs with replacement while keeping corresponding system scores together to estimate uncertainty in a difference.
- Pareto frontier
- Configurations that are not dominated by another measured configuration on the selected objectives.
- pass@k
- The probability that at least one of k generated candidates passes the evaluation tests, estimated under a specified sampling setup.
- Precision
- The fraction of predicted positives that are actually positive: TP / (TP + FP).
- Recall
- The fraction of actual positives that are detected: TP / (TP + FN).
- Retrieval-augmented generation
- Generation that uses documents retrieved for the current request as additional context.
- Rubric
- Explicit scoring criteria, often with anchored examples of different quality levels.
- Slice
- A meaningful subset of evaluation data, such as a language, task, or difficulty level.
- Wilson interval
- A binomial-proportion confidence interval that generally behaves better than a naive normal interval for small samples or extreme rates.
- Agent loop
- A bounded cycle of observing state, choosing an action, executing it, and checking whether to stop.
- Tool contract
- The input schema, authorization, output shape, and error behavior a tool exposes.
- Idempotency
- Repeated requests with the same operation identity have the same intended effect as one request.
- Context window
- The model-specific limit on tokens considered in a request, including required output capacity.
- RAG
- Retrieval-augmented generation: selecting external evidence and providing it as context for an answer.
- Prompt injection
- Untrusted content attempting to redirect a model away from the intended task or policy.
- Trace
- A linked record of the steps, timings, tool calls, and outcomes of one system execution.
- MCP
- Model Context Protocol: a versioned interface for connecting applications to tools and context.
- A2A
- Agent2Agent: a protocol for exchanging tasks, messages, and artifacts across agent systems.
- Checkpoint
- Persisted execution state that supports inspection or resumption after interruption.