MATH (competition mathematics)
Integrity rank #49 of 61 · 79 models scored · top score 98.13 · GPT-5
Competition math; L5=hardest band
Strongest on discrimination, weakest on contamination resistance. Caveat: the public test set does not rule out training-data contamination.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | GPT-5 | 98.13 | reproduced· optimizedT1 | 2025-08-07 |
| 2 | GPT-5 mini | 97.85 | reproduced· optimizedT1 | 2025-08-07 |
| 3 | o4-mini | 97.83 | reproduced· optimizedT1 | 2025-04-16 |
| 4 | o3 | 97.77 | reproduced· optimizedT1 | 2024-12-20 |
| 5 | Claude Sonnet 4.5 | 97.73 | reproduced· optimizedT1 | 2025-09-29 |
| 6 | Qwen3-Max | 97.13 | reproduced· optimizedT1 | 2025-09-05 |
| 7 | DeepSeek-R1 (May 2025) | 96.64 | reproduced· optimizedT1 | 2025-05-28 |
| 8 | o3-mini | 96.49 | reproduced· optimizedT1 | 2025-01-31 |
| 9 | Claude Haiku 4.5 | 96.36 | reproduced· optimizedT1 | 2025-10-15 |
| 10 | Gemini 2.5 Pro (May 2025) | 95.9 | reproduced· optimizedT1 | 2025-05-06 |
| 11 | Gemini 2.5 Pro (Mar 2025) | 95.56 | reproduced· optimizedT1 | 2025-03-25 |
| 12 | GPT-5 nano | 95.24 | reproduced· optimizedT1 | 2025-08-07 |
| 13 | o1 | 94.71 | reproduced· optimizedT1 | 2024-12-05 |
| 14 | DeepSeek-R1 | 93.05 | reproduced· optimizedT1 | 2025-01-20 |
| 15 | Claude 3.7 Sonnet | 91.16 | reproduced· optimizedT1 | 2025-02-24 |
| 16 | Grok-3 mini | 90.94 | reproduced· optimizedT1 | 2025-02-19 |
| 17 | o1-mini | 89.18 | reproduced· optimizedT1 | 2024-09-12 |
| 18 | Grok 3 | 88.75 | reproduced· optimizedT1 | 2025-02-17 |
| 19 | GPT-4.1 mini | 87.29 | reproduced· optimizedT1 | 2025-04-14 |
| 20 | Claude Opus 4 | 85.05 | reproduced· optimizedT1 | 2025-05-22 |
| 21 | Claude Sonnet 4 | 84.37 | reproduced· optimizedT1 | 2025-05-22 |
| 22 | Gemini 2.0 Pro | 83.46 | reproduced· optimizedT1 | 2024-12-11 |
| 23 | GPT-4.1 | 83.01 | reproduced· optimizedT1 | 2025-04-14 |
| 24 | Gemini 2.0 Flash (Feb 2025) | 82.17 | reproduced· optimizedT1 | 2024-12-11 |
| 25 | o1-preview | 81.65 | reproduced· optimizedT1 | 2024-09-12 |
| 26 | Mistral Medium 3 | 81.63 | reproduced· optimizedT1 | 2025-05-07 |
| 27 | GPT-4.5 | 78.63 | reproduced· optimizedT1 | 2025-02-27 |
| 28 | DeepSeek-V3 (Mar 2025) | 75.55 | reproduced· optimizedT1 | 2025-03-24 |
| 29 | Gemma 3 27B | 74.04 | reproduced· optimizedT1 | 2025-03-12 |
| 30 | Llama 4 Maverick | 73.02 | reproduced· optimizedT1 | 2025-04-05 |
| 31 | Gemini 1.5 Pro (Sept 2024) | 70.39 | reproduced· optimizedT1 | 2024-09-24 |
| 32 | GPT-4.1 nano | 70 | reproduced· optimizedT1 | 2025-04-14 |
| 33 | Qwen3-235B-A22B | 68.86 | reproduced· optimizedT1 | 2025-04-28 |
| 34 | Qwen2.5-Max | 67.18 | reproduced· optimizedT1 | 2025-01-25 |
| 35 | Phi-4 | 64.94 | reproduced· optimizedT1 | 2024-12-12 |
| 36 | DeepSeek-V3 | 64.85 | reproduced· optimizedT1 | 2024-12-24 |
| 37 | Grok-2 (Dec 2024) | 63.52 | reproduced· optimizedT1 | 2024-08-13 |
| 38 | Qwen2.5-72B | 63.17 | reproduced· optimizedT1 | 2024-09-19 |
| 39 | Llama 4 Scout | 62.27 | reproduced· optimizedT1 | 2025-04-05 |
| 40 | Gemini 1.5 Flash (Sep 2024) | 61.87 | reproduced· optimizedT1 | 2024-05-10 |
| 41 | Claude 3.5 Sonnet (October 2024) | 56.95 | reproduced· optimizedT1 | 2024-10-22 |
| 42 | GPT-4o (Aug 2024) | 53.28 | reproduced· optimizedT1 | 2024-05-13 |
| 43 | GPT-4o mini | 52.63 | reproduced· optimizedT1 | 2024-07-18 |
| 44 | Claude 3.5 Sonnet | 51.68 | reproduced· optimizedT1 | 2024-06-20 |
| 45 | GPT-4o (May 2024) | 51.05 | reproduced· optimizedT1 | 2024-05-13 |
| 46 | Mistral Large 2 (Nov 2024) | 50.28 | reproduced· optimizedT1 | 2024-07-24 |
| 47 | Llama 3.1-405B | 49.77 | reproduced· optimizedT1 | 2024-07-23 |
| 48 | GPT-4o (Nov 2024) | 49.77 | reproduced· optimizedT1 | 2024-05-13 |
| 49 | Mistral Small 3.1 | 46.77 | reproduced· optimizedT1 | 2025-03-17 |
| 50 | GPT-4 Turbo (Apr 2024) | 46.73 | reproduced· optimizedT1 | 2024-04-09 |
| 51 | Claude 3.5 Haiku | 46.36 | reproduced· optimizedT1 | 2024-10-22 |
| 52 | Mistral Large 2 (Jul 2024) | 44.82 | reproduced· optimizedT1 | 2024-07-24 |
| 53 | Llama 3.3 70B | 41.6 | reproduced· optimizedT1 | 2024-12-06 |
| 54 | Gemini 1.5 Pro (May 2024) | 40.75 | reproduced· optimizedT1 | 2024-05-14 |
| 55 | GPT-4 Turbo (Nov 2023) | 40.02 | reproduced· optimizedT1 | 2023-11-06 |
| 56 | Llama 3.2 90B | 39.44 | reproduced· optimizedT1 | 2024-09-24 |
| 57 | Qwen2-72B | 39.07 | reproduced· optimizedT1 | 2024-06-07 |
| 58 | Claude 3 Opus | 37.48 | reproduced· optimizedT1 | 2024-03-04 |
| 59 | Llama 3.1-70B | 36.68 | reproduced· optimizedT1 | 2024-07-23 |
| 60 | Gemma 2 27B | 27.89 | reproduced· optimizedT1 | 2024-06-24 |
| 61 | Gemini 1.5 Flash (May 2024) | 25.12 | reproduced· optimizedT1 | 2024-05-10 |
| 62 | Mistral Large | 24.46 | reproduced· optimizedT1 | 2024-02-26 |
| 63 | Mixtral 8x22B | 24.24 | reproduced· optimizedT1 | 2024-04-17 |
| 64 | GPT-4 (Jun 2023) | 22.97 | reproduced· optimizedT1 | 2023-06-13 |
| 65 | Llama 3.1-8B | 22.88 | reproduced· optimizedT1 | 2024-07-23 |
| 66 | Llama 3-70B | 22.55 | reproduced· optimizedT1 | 2024-04-18 |
| 67 | Gemma 2 9B | 21.01 | reproduced· optimizedT1 | 2024-06-24 |
| 68 | Claude 3 Sonnet | 18.17 | reproduced· optimizedT1 | 2024-03-04 |
| 69 | phi-3-medium 14B | 17.56 | reproduced· optimizedT1 | 2024-04-23 |
| 70 | GPT-3.5 Turbo (Nov 2023) | 15.89 | reproduced· optimizedT1 | 2023-06-13 |
| 71 | Claude 3 Haiku | 14.88 | reproduced· optimizedT1 | 2024-03-04 |
| 72 | Claude 2 | 11.73 | reproduced· optimizedT1 | 2023-07-11 |
| 73 | GPT-3.5 Turbo (Jan 2024) | 11.63 | reproduced· optimizedT1 | 2023-06-13 |
| 74 | Gemini 1.0 Pro | 11.24 | reproduced· optimizedT1 | 2023-12-06 |
| 75 | Mistral NeMo | 10.83 | reproduced· optimizedT1 | 2024-07-18 |
| 76 | Mixtral 8x7B | 9.95 | reproduced· optimizedT1 | 2023-12-11 |
| 77 | Llama 3-8B | 6.13 | reproduced· optimizedT1 | 2024-04-18 |
| 78 | Yi-34B | 5.15 | reproduced· optimizedT1 | 2023-11-02 |
| 79 | Llama 2-70B | 3.29 | reproduced· optimizedT1 | 2023-07-18 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
This benchmark scores 79 models, producing a 94.84-point spread between the best and worst results. Because the spread is wide, rank differences on this benchmark likely reflect genuine separation between models rather than close, statistically indistinguishable scores.
3 cited facts
This benchmark's Benchmark Integrity Index is 73 (grade B), ranking 49th among 61 benchmarks. The weakest integrity component is contamination, and that narrows how much to trust top scores. As such, the index captures the benchmark's health as a discriminator of frontier models under a disclosed harness, rather than any model's capability.
6 cited facts
This benchmark is saturated. The top score is 98.13, leaving only 1.87 points of headroom to the ceiling. Nine models cluster within a couple of points of the top, so small differences among leaders are close to noise rather than meaningful capability gaps.
5 cited facts
Task performance here comes from a consistent harness, so scores are directly comparable across runs; however, the test set is public, meaning high scores warrant extra scrutiny for possible contamination. This benchmark has a known contamination history—public since 2021 and heavily memorized—so results should be interpreted cautiously. Scores also depend on the disclosed harness conditions: single-shot chain-of-thought, with self-consistency and code tools markedly improving results; these are task-performance figures under that harness, not deployed capability.
4 cited facts
This benchmark measures proficiency in competition mathematics, with difficulty bands culminating in a hardest level. It was built from a large set of competition-archive problems, each tagged by difficulty. Its underlying assumption is that a correct final answer reflects correct reasoning, and that assumption is also the sharpest caveat: guessed answers are rewarded, every problem is public, and code execution tools can trivialize many items, so the benchmark can mislead about genuine reasoning competence.
4 cited facts
MATH level 5 (MATH (competition mathematics)): MATH Level 5 — the hardest (level-5) competition-mathematics problems from the MATH dataset; scored as accuracy (%).
GPT-5 leads MATH level 5 at 98.13 — independently reproduced. The full leaderboard above lists every recorded measurement, not just the headline number.
79 models have recorded scores on MATH level 5, spanning a score spread of 94.84.
tensor.news grades MATH level 5 B for integrity (score 73/100), ranking #49 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
Yes — top scores are clustering near the ceiling (saturation ratio 0.11), so MATH level 5 no longer separates leading models well. Treat small gaps at the top with caution.
Its test set is public, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Scores are reported under a consistent harness, so comparisons on MATH level 5 are reasonably apples-to-apples.
Every MATH level 5 measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites arxiv.org/abs/2103.03874.