Scoring open-ended generation
There can be many good answers to a question, and a fluent answer can still be wrong.
Use the strongest available check
For structured extraction, parse the output and compare normalized fields. For numerical answers, define units and tolerances. For code, execute meaningful tests in an isolated environment. Exact text matching is valuable when the accepted format really is exact; otherwise it can penalize valid alternatives.
Break quality into dimensions
A summary may be faithful, complete, concise, and readable to different degrees. Score those criteria separately with examples of each rating. Word overlap and semantic similarity can be useful signals, but neither proves that a factual statement is supported by the source.
Understand pass@k
For code generation, pass@k estimates whether at least one of k sampled candidates passes the tests. It is not the same as selecting one correct answer without an oracle. Compare results at the same k and sampling setup, and report the cost of generating and testing candidates.
Worked example
Worked example: The reference is “The refund window is 30 days.” An answer saying “You have one month” may or may not be equivalent under the policy. A rubric should decide this explicitly. An answer that copies the reference and invents an exception should fail the unsupported-claim criterion.
Match the grader to the output
Exact match works when only one normalized answer is acceptable. It is brittle when several valid phrasings exist. Token overlap can reward similar wording without establishing factual support. A rubric can distinguish factual correctness, citation support, completeness, and style, but requires clear criteria and calibration.
For generated code, execution against tests is stronger evidence of functional behavior than a judge saying the code looks plausible. The tests still define the measured behavior: incomplete tests can miss bugs, and the execution environment needs resource and security limits.
Understand pass@k
Pass@k measures whether at least one of k generated candidates passes the tests. It is not the probability that the first candidate succeeds, and it does not include the cost or reliability of choosing the good candidate for a user.
When n sampled candidates contain c correct candidates, the common estimator is:
pass@k = 1 − choose(n − c, k) / choose(n, k)
For n = 10, c = 2, and k = 3, the estimate is 1 − 56/120, approximately 53.3%. That does not turn a 20% pass@1 system into a 53.3% single-answer product. Report generation settings, candidate budget, and any selection mechanism.
Your artifact: a rubric with one clear success, one borderline output, and one fluent but unsupported answer.
Key takeaway
Prefer verifiable outcomes, then use explicit rubrics for qualities that need judgment.
Knowledge check
What does pass@10 measure?
- The correctness of the tenth answer
- The chance at least one of ten candidates passes
- A score ten times larger than pass@1
Answer and explanation
The chance at least one of ten candidates passes
pass@k measures success among a set of k candidates. It assumes a way to recognize a passing candidate and uses a larger generation budget than pass@1.
Sources
- Evaluating Large Language Models Trained on Code — Chen et al., 2021. HumanEval, functional correctness, and the pass@k estimator.
Continue learning
- Accuracy, precision, recall & F1 — The same predictions can look excellent or terrible depending on the question your metric asks.
- Scoring open-ended generation — There can be many good answers to a question, and a fluent answer can still be wrong.
- Confidence is not correctness — A model that knows when it might be wrong can be more useful than one that is always certain.