AIME (2024/2025 as LLM benchmark)
Integrity rank #18 of 61 · 8 models scored · top score 79.8 · DeepSeek-R1
HS competition math reasoning, integer answers
Strongest on discrimination, weakest on saturation headroom. Caveat: the public test set does not rule out training-data contamination.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | DeepSeek-R1 | 79.8 | self-reported· optimizedT1 | 2025-01-22 |
| 2 | DeepSeek-R1-Distill-Qwen-32B | 72.6 | self-reported· optimizedT1 | 2025-01-22 |
| 3 | DeepSeek-R1-Distill-Llama-70B | 70 | self-reported· optimizedT1 | 2025-01-22 |
| 4 | DeepSeek-R1-Distill-Qwen-14B | 69.7 | self-reported· optimizedT1 | 2025-01-22 |
| 5 | DeepSeek-R1-Distill-Qwen-7B | 55.5 | self-reported· optimizedT1 | 2025-01-22 |
| 6 | DeepSeek-R1-Distill-Llama-8B | 50.4 | self-reported· optimizedT1 | 2025-01-22 |
| 7 | DeepSeek-V3 | 39.2 | self-reported· optimizedT1 | 2025-01-22 |
| 8 | DeepSeek-R1-Distill-Qwen-1.5B | 28.9 | self-reported· optimizedT1 | 2025-01-22 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
This benchmark scores 8 models, with a score spread of 50.9 points between the highest and lowest results. A spread of this magnitude indicates a wide gap, suggesting that rank differences reflect genuine capability separation rather than noise.
3 cited facts
The integrity score of 97 reflects the benchmark's health as a discriminator of frontier models under a disclosed harness, earning an integrity grade of A. It holds an integrity rank of 15 out of the benchmark universe size, with saturation serving as the weakest component, which means the ceiling is crowded for top performers.
3 cited facts
The benchmark is not currently saturated, featuring a top score of 79.8 and approximately twenty points of headroom to the ceiling. Only one model clusters within a narrow band near the top, highlighting a sparse leading tier. Because the field remains unsaturated, readers should interpret small score differences among the highest performers as genuine capability distinctions rather than statistical noise.
5 cited facts
Scores for this benchmark represent task performance under a disclosed harness, and with harness comparability labeled as consistent harness, they are directly comparable across evaluations, though the reporting method averages over many samples which can inflate stability. Because test set privacy is public, high task performance warrants careful scrutiny for potential data contamination rather than reflecting deployed capability. The benchmark has experienced significant contamination in past iterations, and while recent versions remain cleaner, they remain sensitive to rapid data leakage.
4 cited facts
This benchmark evaluates mathematical reasoning capabilities by requiring models to generate integer solutions for problems sourced from standardized mathematics examinations. The evaluation framework assumes that these contest-style questions effectively measure transferable problem-solving skills and relies on the expectation that recent test versions remain unseen during model training. Practitioners should note that the dataset is exceptionally small and highly variable, meaning individual questions can disproportionately sway results while remaining vulnerable to memorization or accidental exposure.
4 cited facts
AIME (AIME (2024/2025 as LLM benchmark)): AIME (2024/2025 as LLM benchmark) — HS competition math reasoning, integer answers; scored as Accuracy; often avg@k/cons@k.
DeepSeek-R1 leads AIME at 79.8 — self-reported by the lab. The full leaderboard above lists every recorded measurement, not just the headline number.
8 models have recorded scores on AIME, spanning a score spread of 50.9.
tensor.news grades AIME A for integrity (score 97/100), ranking #18 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — AIME still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is public, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Scores are reported under a consistent harness, so comparisons on AIME are reasonably apples-to-apples.
Every AIME measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites matharena.ai.