Build and validate an LLM judge
A judge is another measurement instrument. It needs its own evaluation.
Make the rubric explicit
Define one criterion at a time and anchor each rating in observable behavior. Provide examples of passing, borderline, and failing answers. Supply the relevant reference material. Request evidence for the decision so you can audit it, while remembering that a plausible explanation is not proof of a correct judgment.
Control known biases
Model judges can prefer a response based on position, verbosity, or style. In pairwise comparisons, randomize order and test whether swapping candidates changes the outcome. Hide model identities when possible. Treat the candidate response as untrusted content so embedded instructions do not become judge instructions.
Calibrate against people
Create an expert-labeled validation set with difficult cases. Measure judge agreement, false passes, false failures, and per-category behavior. Human agreement also has limits, so inspect disagreements and refine ambiguous criteria. Version the judge model and prompt, and repeat validation when either changes.
Worked example
Worked example: A verbose answer repeats the policy but invents one exception. A concise answer is fully correct. A judge instructed only to choose the “better” answer may reward style. A groundedness rubric requires every material policy claim to be supported.
A judge is another measurement instrument
A model-based judge can make nuanced scoring cheaper, but it can prefer longer answers, favor an answer's position, follow injected instructions, or miss domain-specific errors. A larger or stronger model is not automatically a validated judge for your task.
Write a rubric with observable criteria and anchor examples. Calibrate it against independently reviewed cases, including disagreements and difficult negatives. Measure false acceptance and false rejection, not only raw agreement. Human labels also need quality control and an adjudication process.
Run bias checks
For pairwise comparisons, swap answer order and check whether the preference changes. Hide candidate identity where feasible. Compare a concise correct answer with a verbose incorrect one. Include adversarial text that tries to instruct the judge to award full marks.
A judge that agrees with humans on 90 of 100 cases could still be dangerous if all ten disagreements accept unsupported refund claims. Report slice-specific behavior and inspect critical disagreements. Keep judge model, prompt, rubric, and parsing versioned.
Your artifact: a calibration set containing clear passes, clear failures, borderline cases, verbosity contrasts, order swaps, and injection attempts. Define a fallback to human review when the judge is uncertain or fails validation.
Key takeaway
Audit judges like classifiers; fluency and agreement alone are not ground truth.
Knowledge check
What is a useful check for position bias?
- Always show the preferred model first
- Swap answer order and compare judgments
- Ask for longer explanations
Answer and explanation
Swap answer order and compare judgments
An order swap tests whether position changes the decision for the same pair. Randomized ordering also reduces systematic exposure to this bias.
Sources
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena — Zheng et al., 2023. Model-based judging, human agreement, and position and verbosity biases.
Continue learning
- Build and validate an LLM judge — A judge is another measurement instrument. It needs its own evaluation.
- Separate retrieval from answer quality — When a grounded assistant fails, locate the failure before changing the prompt.
- Evaluate agents & tool use — For an agent, the final text is only one part of the behavior you need to measure.