Massive Multitask Language Understanding
Integrity rank #53 of 61 · 99 models scored · top score 84.13 · GPT-4o (Nov 2024)
Knowledge breadth across 57 subjects via MC
Strongest on discrimination, weakest on contamination resistance. Caveat: scores come from an inconsistent mix of harnesses, so head-to-head comparisons are not apples-to-apples.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | GPT-4o (Nov 2024) | 84.13 | self-reported· optimizedT1 | 2024-05-13 |
| 2 | Claude 3.5 Sonnet (October 2024) | 83.07 | unverified· optimizedT1 | 2024-10-22 |
| 3 | DeepSeek-V3 | 82.93 | unverified· optimizedT1 | 2024-12-24 |
| 4 | Gemini 1.5 Pro (Sept 2024) | 82.53 | unverified· optimizedT1 | 2024-09-24 |
| 5 | Claude 3.5 Sonnet | 82 | unverified· optimizedT1 | 2024-06-20 |
| 6 | GPT-4 (Mar 2023) | 81.87 | self-reported· optimizedT1 | 2023-03-15 |
| 7 | Llama 3.3 70B | 81.73 | unverified· optimizedT1 | 2024-12-06 |
| 8 | Gemini 1.5 Pro (May 2024) | 81.2 | unverified· optimizedT1 | 2024-05-14 |
| 9 | Qwen2.5-72B | 80.4 | self-reported· optimizedT1 | 2024-09-19 |
| 10 | Phi-4 | 79.73 | self-reported· optimizedT1 | 2024-12-12 |
| 11 | Claude 3 Opus | 79.47 | unverified· optimizedT1 | 2024-03-04 |
| 12 | Llama 3.1-405B | 79.33 | unverified· optimizedT1 | 2024-07-23 |
| 13 | GPT-4o (Aug 2024) | 79.07 | unverified· optimizedT1 | 2024-05-13 |
| 14 | GPT-4o (May 2024) | 78.93 | unverified· optimizedT1 | 2024-05-13 |
| 15 | Qwen2-72B | 76.53 | unverified· optimizedT1 | 2024-06-07 |
| 16 | GPT-4 (Jun 2023) | 76.53 | unverified· optimizedT1 | 2023-06-13 |
| 17 | Amazon Nova Pro | 76 | unverified· optimizedT1 | 2024-12-03 |
| 18 | GPT-4o mini | 75.73 | self-reported· optimizedT1 | 2024-07-18 |
| 19 | GPT-4 Turbo (Apr 2024) | 75.07 | unverified· optimizedT1 | 2024-04-09 |
| 20 | Llama 3.2 90B | 73.73 | unverified· optimizedT1 | 2024-09-24 |
| 21 | Llama 3.1-70B | 73.47 | unverified· optimizedT1 | 2024-07-23 |
| 22 | Mistral Large 2 (Jul 2024) | 73.33 | unverified· optimizedT1 | 2024-07-24 |
| 23 | Gemini 2.0 Flash (Dec 2024) | 72.93 | unverified· optimizedT1 | 2024-12-11 |
| 24 | GPT-4 Turbo (Nov 2023) | 72.8 | unverified· optimizedT1 | 2023-11-06 |
| 25 | Llama 3-70B | 72.4 | unverified· optimizedT1 | 2024-04-18 |
| 26 | Qwen2.5-Coder-32B | 72.13 | self-reported· optimizedT1 | 2024-09-18 |
| 27 | Claude 2 | 71.33 | self-reported· optimizedT1 | 2023-07-11 |
| 28 | DeepSeek-V2 (MoE-236B, May 2024) | 71.2 | self-reported· optimizedT1 | 2024-05-07 |
| 29 | phi-3-medium 14B | 70.67 | unverified· optimizedT1 | 2024-04-23 |
| 30 | Gemini 1.5 Flash (May 2024) | 70.53 | unverified· optimizedT1 | 2024-05-10 |
| 31 | Mixtral 8x22B | 70.4 | unverified· optimizedT1 | 2024-04-17 |
| 32 | Yi-34B | 68.4 | unverified· optimizedT1 | 2023-11-02 |
| 33 | Claude 3 Sonnet | 67.87 | unverified· optimizedT1 | 2024-03-04 |
| 34 | Gemma 2 27B | 67.6 | unverified· optimizedT1 | 2024-06-24 |
| 35 | phi-3-small 7.4B | 67.6 | self-reported· optimizedT1 | 2024-04-23 |
| 36 | Qwen2.5-Coder-14B | 66.93 | self-reported· optimizedT1 | 2024-09-18 |
| 37 | Claude 3.5 Haiku | 65.73 | unverified· optimizedT1 | 2024-10-22 |
| 38 | Gemini 1.5 Flash (Sep 2024) | 65.2 | unverified· optimizedT1 | 2024-05-10 |
| 39 | Claude 3 Haiku | 65.07 | unverified· optimizedT1 | 2024-03-04 |
| 40 | Claude 2.1 | 64.67 | unverified· optimizedT1 | 2023-11-21 |
| 41 | Claude Instant | 64.53 | self-reported· optimizedT1 | 2023-08-09 |
| 42 | Gemma 2 9B | 62.8 | unverified· optimizedT1 | 2024-06-24 |
| 43 | GPT-3.5 Turbo (Nov 2023) | 61.87 | self-reported· optimizedT1 | 2023-06-13 |
| 44 | Mixtral 8x7B | 60.8 | unverified· optimizedT1 | 2023-12-11 |
| 45 | Falcon-180B | 60.8 | unverified· optimizedT1 | 2023-09-06 |
| 46 | Gemini 1.0 Pro | 60 | unverified· optimizedT1 | 2023-12-06 |
| 47 | Llama 2-70B | 59.87 | unverified· optimizedT1 | 2023-07-18 |
| 48 | GPT-3.5 Turbo (Jun 2023) | 58.53 | unverified· optimizedT1 | 2023-06-13 |
| 49 | Llama 3-8B | 58.4 | unverified· optimizedT1 | 2024-04-18 |
| 50 | Mistral Large | 58.4 | unverified· optimizedT1 | 2024-02-26 |
| 51 | phi-3-mini 3.8B | 58.4 | self-reported· optimizedT1 | 2024-04-23 |
| 52 | Stable Beluga 2 | 58.13 | self-reported· optimizedT1 | 2023-07-20 |
| 53 | Yi-9B | 57.87 | unverified· optimizedT1 | 2024-03-01 |
| 54 | Qwen2.5-Coder (7B) | 57.33 | self-reported· optimizedT1 | 2024-09-18 |
| 55 | GPT-3.5 Turbo (Jan 2024) | 56.4 | unverified· optimizedT1 | 2023-06-13 |
| 56 | Qwen-14B | 55.07 | self-reported· optimizedT1 | 2023-09-24 |
| 57 | Gemma 7B | 54.8 | unverified· optimizedT1 | 2024-02-21 |
| 58 | StarCoder 2 15B | 52.13 | self-reported· optimizedT1 | 2024-02-29 |
| 59 | Yi 6B | 52 | unverified· optimizedT1 | 2023-11-02 |
| 60 | LLaMA-65B | 51.2 | unverified· optimizedT1 | 2023-02-24 |
| 61 | Llama 2-34B | 50.13 | unverified· optimizedT1 | 2023-07-18 |
| 62 | Mistral 7B v0.1 | 50 | unverified· optimizedT1 | 2023-10-10 |
| 63 | DeepSeek-Coder-V2-Lite-Base | 47.33 | self-reported· optimizedT1 | 2024-06-13 |
| 64 | Baichuan2-13B | 45.6 | unverified· optimizedT1 | 2023-09-06 |
| 65 | LLaMA-33B | 44.93 | unverified· optimizedT1 | 2023-02-27 |
| 66 | Nemotron-4 15B | 44.93 | self-reported· optimizedT1 | 2024-02-27 |
| 67 | Phi-2 | 44.53 | unverified· optimizedT1 | 2023-12-12 |
| 68 | Falcon 2 11B | 44.53 | self-reported· optimizedT1 | 2024-05-09 |
| 69 | Falcon-40B | 42.53 | self-reported· optimizedT1 | 2023-03-15 |
| 70 | Llama 3.1-8B | 41.47 | unverified· optimizedT1 | 2024-07-23 |
| 71 | Llama 2-13B | 40.8 | unverified· optimizedT1 | 2023-07-18 |
| 72 | Baichuan 2-7B | 38.88 | unverified· optimizedT1 | 2023-09-20 |
| 73 | Qwen2.5-Coder (1.5B) | 38.13 | self-reported· optimizedT1 | 2024-09-18 |
| 74 | internlm-7b | 34.67 | self-reported· optimizedT1 | 2023-07-05 |
| 75 | INTELLECT-1 | 33.19 | self-reported· optimizedT1 | 2024-11-29 |
| 76 | MPT-30B | 30.53 | self-reported· optimizedT1 | 2023-06-22 |
| 77 | chatglm2-6b | 30.53 | self-reported· optimizedT1 | 2023-06-24 |
| 78 | LLaMA-13B | 30.27 | unverified· optimizedT1 | 2023-02-27 |
| 79 | Llama 2-7B | 27.73 | unverified· optimizedT1 | 2023-07-18 |
| 80 | Qwen-7B | 26.67 | self-reported· optimizedT1 | 2023-09-28 |
| 81 | Gemma 2B | 23.07 | unverified· optimizedT1 | 2024-02-21 |
| 82 | Baichuan1-7B | 23.07 | unverified· optimizedT1 | 2023-06-01 |
| 83 | Qwen2.5-Coder-0.5B | 22.67 | self-reported· optimizedT1 | 2024-09-18 |
| 84 | CodeQwen1.5-7B | 20.67 | self-reported· optimizedT1 | 2024-04-15 |
| 85 | DeepSeek Coder 33B | 19.2 | self-reported· optimizedT1 | 2024-01-25 |
| 86 | StarCoder 2 7B | 18.4 | self-reported· optimizedT1 | 2024-02-29 |
| 87 | Phi-1.5 | 16.8 | self-reported· optimizedT1 | 2023-09-11 |
| 88 | StarCoder 2 3B | 15.47 | self-reported· optimizedT1 | 2024-02-29 |
| 89 | DeepSeek Coder 6.7B | 15.2 | self-reported· optimizedT1 | 2024-01-25 |
| 90 | XGen-7B | 15.07 | self-reported· optimizedT1 | 2023-09-07 |
| 91 | LLaMA-7B | 14.13 | unverified· optimizedT1 | 2023-02-24 |
| 92 | Falcon-7B | 13.33 | self-reported· optimizedT1 | 2023-04-24 |
| 93 | MPT-7B | 7.73 | self-reported· optimizedT1 | 2023-05-05 |
| 94 | open_llama_7b | 6.53 | self-reported· optimizedT1 | 2023-06-07 |
| 95 | Qwen-1_8B | 4.27 | self-reported· optimizedT1 | 2023-11-30 |
| 96 | RedPajama-INCITE-7B-Base | 1.73 | self-reported· optimizedT1 | 2023-05-04 |
| 97 | Dolly 2.0-12b | 1.6 | self-reported· optimizedT1 | 2023-04-12 |
| 98 | Cerebras-GPT-13B | 1.6 | self-reported· optimizedT1 | 2023-04-06 |
| 99 | DeepSeek Coder 1.3B | 1.07 | self-reported· optimizedT1 | 2024-01-25 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
The benchmark's Benchmark Integrity Index is 67 (grade C), ranking 53rd out of 61 benchmarks. This index reflects the benchmark's health as a discriminator of frontier models under a disclosed harness; contamination is the weakest integrity component, and its presence narrows how much to trust top scores.
6 cited facts
The benchmark is not saturated: the top score is 84.13, the remaining headroom to the ceiling is 15.87 points, and four models cluster within a couple of points near the top. Because the benchmark is not saturated, the ranking retains genuine separation at the top, so a gap of a few points reflects a real capability difference rather than noise.
5 cited facts
This benchmark's scores are produced under a mixed harness and are not directly comparable, so task performance measured by different harnesses should not be compared head-to-head. Because the test set is public, high scores warrant more scrutiny for possible contamination than a held-out test set would. This benchmark has been treated as contaminated, and answer-extraction and format choices introduce notable sensitivity in measured performance; these scores are task performance under a disclosed harness, not deployed capability.
4 cited facts
This benchmark assesses breadth of knowledge across dozens of subjects through multiple-choice items, on the assumption that this format can stand in for general knowledge. Its items were collected by students from exams and textbooks, giving it a broad but uneven source base. The sharpest caveat is that a nontrivial fraction of items are mislabeled or ambiguous, and because the set is heavily contaminated, scores should not be read as a clean measure of a knowledge ceiling.
4 cited facts
MMLU (Massive Multitask Language Understanding): Massive Multitask Language Understanding — a 57-subject multiple-choice knowledge and reasoning benchmark; scored as accuracy (%). Widely regarded as saturated for frontier models.
GPT-4o (Nov 2024) leads MMLU at 84.13 — self-reported by the lab. The full leaderboard above lists every recorded measurement, not just the headline number.
99 models have recorded scores on MMLU, spanning a score spread of 83.06.
tensor.news grades MMLU C for integrity (score 67/100), ranking #53 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — MMLU still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is public, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Only partly — scores come from a mix of harnesses and protocols, so a head-to-head on MMLU is not fully apples-to-apples.
Every MMLU measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites arxiv.org/abs/2009.03300.