# What is an LLM evaluation?

Canonical URL: https://agentlearn.dev/learn/evals/what-is-eval
Author: [Hemanth HM](https://h3manth.com)
Track: evals
Reading time: 8 minutes

An impressive answer is an observation. An evaluation turns many observations into evidence for a decision.

## Start with a decision

Imagine a support assistant that answers questions about returns. “Is this a good model?” is too broad to test. “Can this assistant answer return-policy questions accurately, with a supporting citation?” identifies a task, an expected behavior, and a user need. Write that claim before collecting examples.

## The four parts of an eval

An evaluation needs inputs, a system under test, a scoring rule, and an analysis. The system includes the model, prompt, retrieval, tools, and generation settings. A score only has meaning relative to that setup. A benchmark standardizes some of these pieces so different systems can be compared.

## A score is a measurement

Your examples are a sample of possible interactions. A high pass rate may hide failures on a small but important category. Keep the individual outputs, inspect errors, and report which population your sample represents. Every score should come with a task description, sample size, and known limitations.

## Worked example

Worked example: On 100 fictional support questions, 82 answers match the policy and cite the correct paragraph. The observed joint pass rate is 82%. This does not establish an 82% success rate for every future user or prove that the other 18 answers share the same failure.

## Write an evaluation contract

For the support assistant, write the claim as: “Given a current policy passage and a customer's question, the assistant provides a supported answer or explicitly abstains.” Then specify the population: English-language returns questions for the shop, excluding payment execution. This prevents a result on one narrow task from quietly becoming a claim about all customer support.

A case should have a stable ID, input, expected behavior, relevant evidence, slice labels, and scoring guidance. A run should record the case ID, full system version, output, tool trace, grader version, latency, and cost. Store sensitive fields only when necessary and under an explicit access and retention policy.

## Read the denominator

Suppose 80 of 100 requests receive a correct answer, 10 correctly abstain, and 10 fail. “Accuracy” could mean 80% if abstentions do not count as answers, or 90% if the task is correct answer-or-abstention behavior. Neither number is meaningful without the scoring definition. Report answer coverage separately from correctness among answered cases.

**Your artifact:** write a one-paragraph evaluation contract and three examples: ordinary, ambiguous, and unanswerable. Another person should be able to score them without asking what “good” means.


## Key takeaway

An eval is a repeatable test of a specific claim about a system.

## Knowledge check

Which is the most testable evaluation objective?

1. Find the smartest language model
2. Measure correct, cited answers on held-out return-policy questions
3. Get a model to sound confident

Answer: Measure correct, cited answers on held-out return-policy questions

The second objective specifies the task, success criteria, and a separate test set. “Smartest” and “confident” do not define the user outcome.

## Sources

- [Holistic Evaluation of Language Models](https://arxiv.org/abs/2211.09110) — Liang et al., 2022. A framework for evaluating multiple dimensions of model behavior, beyond one headline score.
