Benchmark explainer
What is MMLU?
Knowledge breadth across 57 subjects via MC
Massive Multitask Language Understanding — a 57-subject multiple-choice knowledge and reasoning benchmark; scored as accuracy (%). Widely regarded as saturated for frontier models.
How it's scored
- Metric
- MC accuracy (4-option)
- Score ceiling
- 100
- Construction
- ~15.9k Qs scraped from exams/textbooks by students
- Human baseline
- ~89.8% for 95th-pctile specialist
How trustworthy
Discrimination
100/100Does it still separate models?
Saturation headroom
96.6/100How far from ceiling / clustered at the top?
Contamination resistance
0/100Public vs held-out; training-leak risk.
Harness comparability
50/100Is it apples-to-apples, or mixed / vendor-optimized?
Freshness
0.7/100How old is the benchmark?
Contamination history: Fully public since 2020, ubiquitous in training; treated as contaminated
Limitations: ~6.5% mislabeled/ambiguous (MMLU-Redux) caps meaningful score <100%; pure recall; heavily contaminated
Who leads MMLU
| Model | Score | Evidence |
|---|---|---|
| GPT-4o (Nov 2024) | 84.13 | self-reported |
| Claude 3.5 Sonnet (October 2024) | 83.07 | unverified |
| DeepSeek-V3 | 82.93 | unverified |
| Gemini 1.5 Pro (Sept 2024) | 82.53 | unverified |
| Claude 3.5 Sonnet | 82 | unverified |
Frequently asked questions
MMLU (Massive Multitask Language Understanding): Massive Multitask Language Understanding — a 57-subject multiple-choice knowledge and reasoning benchmark; scored as accuracy (%). Widely regarded as saturated for frontier models.
GPT-4o (Nov 2024) leads MMLU at 84.13 — self-reported by the lab. The full leaderboard above lists every recorded measurement, not just the headline number.
99 models have recorded scores on MMLU, spanning a score spread of 83.06.
tensor.news grades MMLU C for integrity (score 67/100), ranking #53 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — MMLU still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is public, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Only partly — scores come from a mix of harnesses and protocols, so a head-to-head on MMLU is not fully apples-to-apples.
Every MMLU measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites arxiv.org/abs/2009.03300.