Skip to content

Benchmark explainer

What is MMLU?

Knowledge breadth across 57 subjects via MC

Massive Multitask Language Understanding — a 57-subject multiple-choice knowledge and reasoning benchmark; scored as accuracy (%). Widely regarded as saturated for frontier models.

CINTEGRITY 67 / 100tensor.news
mixed harness — not comparablepublic test set

How it's scored

Metric
MC accuracy (4-option)
Score ceiling
100
Construction
~15.9k Qs scraped from exams/textbooks by students
Human baseline
~89.8% for 95th-pctile specialist

How trustworthy

Discrimination

100/100

Does it still separate models?

Saturation headroom

96.6/100

How far from ceiling / clustered at the top?

Contamination resistance

0/100

Public vs held-out; training-leak risk.

Harness comparability

50/100

Is it apples-to-apples, or mixed / vendor-optimized?

Freshness

0.7/100

How old is the benchmark?

Contamination history: Fully public since 2020, ubiquitous in training; treated as contaminated

Limitations: ~6.5% mislabeled/ambiguous (MMLU-Redux) caps meaningful score <100%; pure recall; heavily contaminated

Who leads MMLU

ModelScoreEvidence
GPT-4o (Nov 2024)84.13self-reported
Claude 3.5 Sonnet (October 2024)83.07unverified
DeepSeek-V382.93unverified
Gemini 1.5 Pro (Sept 2024)82.53unverified
Claude 3.5 Sonnet82unverified

Frequently asked questions

MMLU (Massive Multitask Language Understanding): Massive Multitask Language Understanding — a 57-subject multiple-choice knowledge and reasoning benchmark; scored as accuracy (%). Widely regarded as saturated for frontier models.

GPT-4o (Nov 2024) leads MMLU at 84.13 — self-reported by the lab. The full leaderboard above lists every recorded measurement, not just the headline number.

99 models have recorded scores on MMLU, spanning a score spread of 83.06.

tensor.news grades MMLU C for integrity (score 67/100), ranking #53 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.

No — MMLU still has headroom and continues to discriminate between models rather than bunching them at the ceiling.

Its test set is public, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.

Only partly — scores come from a mix of harnesses and protocols, so a head-to-head on MMLU is not fully apples-to-apples.

Every MMLU measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites arxiv.org/abs/2009.03300.

Follow the record