AgentLearn

AI Agents & LLM Evaluation Glossary

Understand key AI agent and evaluation terms, from precision, recall, and calibration to RAG, tool contracts, prompt injection, MCP, and agent traces.

Ablation
An experiment that removes or changes one component to understand its contribution.
Accuracy
The fraction of predictions that are correct: (TP + TN) / all examples.
Benchmark
A standardized dataset and evaluation protocol for comparing systems on a defined task.
Calibration
Agreement between predicted probabilities and observed frequencies of correctness.
Confidence interval
An interval from a procedure designed to cover a fixed population parameter at a stated long-run rate.
Contamination
Exposure to test questions, answers, or close equivalents during training or development.
Coverage
The fraction of requests a system chooses to answer rather than abstain or defer.
Distribution shift
A change in the input or outcome distribution between evaluation and deployment.
F1 score
The harmonic mean of precision and recall: 2PR / (P + R). It does not account for true negatives.
Groundedness
The extent to which an answer’s claims are supported by the provided evidence.
Holdout
Data kept separate from development choices for a less biased evaluation of the finished system.
LLM-as-a-judge
A language model used to apply a rubric or compare outputs; the judge needs independent validation.
Non-inferiority
Testing whether a candidate is no worse than a baseline by more than a predefined meaningful margin.
Paired bootstrap
Resampling inputs with replacement while keeping corresponding system scores together to estimate uncertainty in a difference.
Pareto frontier
Configurations that are not dominated by another measured configuration on the selected objectives.
pass@k
The probability that at least one of k generated candidates passes the evaluation tests, estimated under a specified sampling setup.
Precision
The fraction of predicted positives that are actually positive: TP / (TP + FP).
Recall
The fraction of actual positives that are detected: TP / (TP + FN).
Retrieval-augmented generation
Generation that uses documents retrieved for the current request as additional context.
Rubric
Explicit scoring criteria, often with anchored examples of different quality levels.
Slice
A meaningful subset of evaluation data, such as a language, task, or difficulty level.
Wilson interval
A binomial-proportion confidence interval that generally behaves better than a naive normal interval for small samples or extreme rates.
Agent loop
A bounded cycle of observing state, choosing an action, executing it, and checking whether to stop.
Tool contract
The input schema, authorization, output shape, and error behavior a tool exposes.
Idempotency
Repeated requests with the same operation identity have the same intended effect as one request.
Context window
The model-specific limit on tokens considered in a request, including required output capacity.
RAG
Retrieval-augmented generation: selecting external evidence and providing it as context for an answer.
Prompt injection
Untrusted content attempting to redirect a model away from the intended task or policy.
Trace
A linked record of the steps, timings, tool calls, and outcomes of one system execution.
MCP
Model Context Protocol: a versioned interface for connecting applications to tools and context.
A2A
Agent2Agent: a protocol for exchanging tasks, messages, and artifacts across agent systems.
Checkpoint
Persisted execution state that supports inspection or resumption after interruption.