What is an LLM evaluation?
An impressive answer is an observation. An evaluation turns many observations into evidence for a decision.
Start with a decision
Imagine a support assistant that answers questions about returns. “Is this a good model?” is too broad to test. “Can this assistant answer return-policy questions accurately, with a supporting citation?” identifies a task, an expected behavior, and a user need. Write that claim before collecting examples.
The four parts of an eval
An evaluation needs inputs, a system under test, a scoring rule, and an analysis. The system includes the model, prompt, retrieval, tools, and generation settings. A score only has meaning relative to that setup. A benchmark standardizes some of these pieces so different systems can be compared.
A score is a measurement
Your examples are a sample of possible interactions. A high pass rate may hide failures on a small but important category. Keep the individual outputs, inspect errors, and report which population your sample represents. Every score should come with a task description, sample size, and known limitations.
Worked example
Worked example: On 100 fictional support questions, 82 answers match the policy and cite the correct paragraph. The observed joint pass rate is 82%. This does not establish an 82% success rate for every future user or prove that the other 18 answers share the same failure.
Write an evaluation contract
For the support assistant, write the claim as: “Given a current policy passage and a customer's question, the assistant provides a supported answer or explicitly abstains.” Then specify the population: English-language returns questions for the shop, excluding payment execution. This prevents a result on one narrow task from quietly becoming a claim about all customer support.
A case should have a stable ID, input, expected behavior, relevant evidence, slice labels, and scoring guidance. A run should record the case ID, full system version, output, tool trace, grader version, latency, and cost. Store sensitive fields only when necessary and under an explicit access and retention policy.
Read the denominator
Suppose 80 of 100 requests receive a correct answer, 10 correctly abstain, and 10 fail. “Accuracy” could mean 80% if abstentions do not count as answers, or 90% if the task is correct answer-or-abstention behavior. Neither number is meaningful without the scoring definition. Report answer coverage separately from correctness among answered cases.
Your artifact: write a one-paragraph evaluation contract and three examples: ordinary, ambiguous, and unanswerable. Another person should be able to score them without asking what “good” means.
Key takeaway
An eval is a repeatable test of a specific claim about a system.
Knowledge check
Which is the most testable evaluation objective?
- Find the smartest language model
- Measure correct, cited answers on held-out return-policy questions
- Get a model to sound confident
Answer and explanation
Measure correct, cited answers on held-out return-policy questions
The second objective specifies the task, success criteria, and a separate test set. “Smartest” and “confident” do not define the user outcome.
Sources
- Holistic Evaluation of Language Models — Liang et al., 2022. A framework for evaluating multiple dimensions of model behavior, beyond one headline score.
Continue learning
- What is an LLM evaluation? — An impressive answer is an observation. An evaluation turns many observations into evidence for a decision.
- Build a representative dataset — What you choose to test determines what you are able to discover.
- Baselines, controls & reproducibility — A new score becomes useful when you can explain what changed and what it improved upon.