LLM Benchmark Explorer: Tasks & Limitations
Understand MMLU, GSM8K, HumanEval, SWE-bench, GPQA, and MT-Bench: what each measures, its limitations, and links to official benchmarks and papers.
These are benchmark descriptions, not a live leaderboard. Lab measurements use synthetic teaching data.
MMLU
Academic and professional knowledge, from history to computer science.
Metric: Multiple-choice accuracy. Scope: 57 subjects.
Limitation: Sensitive to prompt format; does not measure grounded conversation.
- Measuring Massive Multitask Language Understanding — Hendrycks et al., 2020. The original MMLU paper: multiple-choice evaluation across 57 subjects.
GSM8K
Grade-school math problems that require multiple reasoning steps.
Metric: Final-answer accuracy. Scope: 8.5k problems · full dataset.
Limitation: Answer extraction matters; a correct result can conceal faulty reasoning.
- Training Verifiers to Solve Math Word Problems — Cobbe et al., 2021. GSM8K and the use of verifiers for multi-step mathematical problem solving.
HumanEval
Python function completion, evaluated with executable unit tests.
Metric: pass@k. Scope: 164 programming tasks.
Limitation: Small, public dataset; test coverage and sampling budget affect interpretation.
- Evaluating Large Language Models Trained on Code — Chen et al., 2021. HumanEval, functional correctness, and the pass@k estimator.
SWE-bench
Real GitHub issues that require changes across a Python repository.
Metric: Percentage of issues resolved. Scope: 2,294 issues · original dataset.
Limitation: Original, Lite, and Verified are different sets. Agent setup and budget matter.
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues? — Jimenez et al., 2023. Repository-level software engineering evaluation using real issues and executable tests.
GPQA
Challenging expert-written questions in biology, physics, and chemistry.
Metric: Multiple-choice accuracy. Scope: 448 questions · original dataset.
Limitation: The Diamond subset differs from the full set; expert labels can still be ambiguous.
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark — Rein et al., 2023. Expert-written questions in biology, physics, and chemistry.
MT-Bench
Open-ended, multi-turn conversations evaluated using a model judge.
Metric: Judge ratings. Scope: 80 multi-turn questions.
Limitation: Judge version, position, and verbosity biases can affect ratings.
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena — Zheng et al., 2023. Model-based judging, human agreement, and position and verbosity biases.