MMLU, GSM8K, HumanEval & beyond
A benchmark is a lens on a capability, not a universal intelligence score.
Map the task before the score
MMLU measures multiple-choice performance across 57 academic and professional subjects. It is useful for breadth of knowledge, but its question format differs from a multi-turn assistant workflow. Ask which behaviors your application shares with the benchmark and which it does not.
Use complementary task families
GSM8K focuses on grade-school mathematical word problems. HumanEval uses Python function-completion tasks with tests. SWE-bench asks systems to resolve repository issues. These tasks stress different combinations of reasoning, code generation, tools, and environment interaction; their percentages do not share a common difficulty scale.
Read the evaluation protocol
Identify the exact dataset version, split, few-shot setup, answer extraction, tool access, and number of candidates. Check whether a score includes retries or a verifier. A leaderboard row is the output of a protocol, and a protocol mismatch can make two apparently identical scores incomparable.
Worked example
Worked example: A model with 86% multiple-choice accuracy may still cite nonexistent documents in a support workflow. Use the public result to form a hypothesis, then build an application-specific grounded-answer evaluation to test it.
Read the benchmark's measurement contract
A benchmark score compresses tasks, inputs, scoring, and execution rules. MMLU covers multiple-choice knowledge questions across 57 subjects; HumanEval tests generated programs against executable tests; SWE-bench evaluates repository-level issue resolution. These measure different activities, so their percentages are not interchangeable units of general intelligence.
For any result, record the benchmark version and subset, prompting method, number of attempts, tool access, inference budget, grader, and date. A changed harness can change the score without a changed model. A subset result should not be labeled as the entire benchmark.
Transfer carefully to your product
A coding benchmark may help select candidates for a coding assistant. It does not establish whether a support assistant cites the current return policy or respects customer ownership. Use public benchmarks as background evidence and your own task evaluation as the release instrument.
The benchmark explorer intentionally describes tasks and limitations instead of presenting a live leaderboard. A static score table would become stale and could imply comparability across incompatible setups.
Your artifact: pick one benchmark and write two statements: “This gives evidence about…” and “This does not establish…”. Then identify one product-specific evaluation that fills the gap.
Key takeaway
Choose benchmarks by construct and protocol, then validate on your own task.
Knowledge check
Which benchmark most directly tests repository issue resolution?
- MMLU
- SWE-bench
- GSM8K
Answer and explanation
SWE-bench
SWE-bench evaluates code changes for real repository issues. MMLU tests subject knowledge and GSM8K tests math word problems.
Sources
- Measuring Massive Multitask Language Understanding — Hendrycks et al., 2020. The original MMLU paper: multiple-choice evaluation across 57 subjects.
Continue learning
- MMLU, GSM8K, HumanEval & beyond — A benchmark is a lens on a capability, not a universal intelligence score.
- Contamination, saturation & leakage — A system that has seen the answers may look capable without demonstrating generalization.
- Quality, latency & cost tradeoffs — The best system is often the one that satisfies the task at an acceptable operating cost.