AgentLearn

MMLU, GSM8K, HumanEval & beyond

A benchmark is a lens on a capability, not a universal intelligence score.

Map the task before the score

MMLU measures multiple-choice performance across 57 academic and professional subjects. It is useful for breadth of knowledge, but its question format differs from a multi-turn assistant workflow. Ask which behaviors your application shares with the benchmark and which it does not.

Use complementary task families

GSM8K focuses on grade-school mathematical word problems. HumanEval uses Python function-completion tasks with tests. SWE-bench asks systems to resolve repository issues. These tasks stress different combinations of reasoning, code generation, tools, and environment interaction; their percentages do not share a common difficulty scale.

Read the evaluation protocol

Identify the exact dataset version, split, few-shot setup, answer extraction, tool access, and number of candidates. Check whether a score includes retries or a verifier. A leaderboard row is the output of a protocol, and a protocol mismatch can make two apparently identical scores incomparable.

Worked example

Worked example: A model with 86% multiple-choice accuracy may still cite nonexistent documents in a support workflow. Use the public result to form a hypothesis, then build an application-specific grounded-answer evaluation to test it.

Read the benchmark's measurement contract

A benchmark score compresses tasks, inputs, scoring, and execution rules. MMLU covers multiple-choice knowledge questions across 57 subjects; HumanEval tests generated programs against executable tests; SWE-bench evaluates repository-level issue resolution. These measure different activities, so their percentages are not interchangeable units of general intelligence.

For any result, record the benchmark version and subset, prompting method, number of attempts, tool access, inference budget, grader, and date. A changed harness can change the score without a changed model. A subset result should not be labeled as the entire benchmark.

Transfer carefully to your product

A coding benchmark may help select candidates for a coding assistant. It does not establish whether a support assistant cites the current return policy or respects customer ownership. Use public benchmarks as background evidence and your own task evaluation as the release instrument.

The benchmark explorer intentionally describes tasks and limitations instead of presenting a live leaderboard. A static score table would become stale and could imply comparability across incompatible setups.

Your artifact: pick one benchmark and write two statements: “This gives evidence about…” and “This does not establish…”. Then identify one product-specific evaluation that fills the gap.

Key takeaway

Choose benchmarks by construct and protocol, then validate on your own task.

Knowledge check

Which benchmark most directly tests repository issue resolution?

  1. MMLU
  2. SWE-bench
  3. GSM8K
Answer and explanation

SWE-bench

SWE-bench evaluates code changes for real repository issues. MMLU tests subject knowledge and GSM8K tests math word problems.

Sources

Read this lesson as Markdown

Continue learning