AgentLearn

LLM Benchmark Explorer: Tasks & Limitations

Understand MMLU, GSM8K, HumanEval, SWE-bench, GPQA, and MT-Bench: what each measures, its limitations, and links to official benchmarks and papers.

These are benchmark descriptions, not a live leaderboard. Lab measurements use synthetic teaching data.

MMLU

Academic and professional knowledge, from history to computer science.

Metric: Multiple-choice accuracy. Scope: 57 subjects.

Limitation: Sensitive to prompt format; does not measure grounded conversation.

GSM8K

Grade-school math problems that require multiple reasoning steps.

Metric: Final-answer accuracy. Scope: 8.5k problems · full dataset.

Limitation: Answer extraction matters; a correct result can conceal faulty reasoning.

HumanEval

Python function completion, evaluated with executable unit tests.

Metric: pass@k. Scope: 164 programming tasks.

Limitation: Small, public dataset; test coverage and sampling budget affect interpretation.

SWE-bench

Real GitHub issues that require changes across a Python repository.

Metric: Percentage of issues resolved. Scope: 2,294 issues · original dataset.

Limitation: Original, Lite, and Verified are different sets. Agent setup and budget matter.

GPQA

Challenging expert-written questions in biology, physics, and chemistry.

Metric: Multiple-choice accuracy. Scope: 448 questions · original dataset.

Limitation: The Diamond subset differs from the full set; expert labels can still be ambiguous.

MT-Bench

Open-ended, multi-turn conversations evaluated using a model judge.

Metric: Judge ratings. Scope: 80 multi-turn questions.

Limitation: Judge version, position, and verbosity biases can affect ratings.