Benchmark explainer
What is AIME?
HS competition math reasoning, integer answers
AIME (2024/2025 as LLM benchmark) — HS competition math reasoning, integer answers; scored as Accuracy; often avg@k/cons@k.
How it's scored
- Metric
- Accuracy; often avg@k/cons@k
- Score ceiling
- 100
- Construction
- MAA exam; 15 Qs/exam, 2 exams/yr; integer answers 0-999
- Human baseline
- AIME qualifiers avg ~5-8/15; ~10+ USAMO-qualifying
How trustworthy
Discrimination
100/100Does it still separate models?
Saturation headroom
90/100How far from ceiling / clustered at the top?
Contamination resistance
100/100Public vs held-out; training-leak risk.
Harness comparability
100/100Is it apples-to-apples, or mixed / vendor-optimized?
Freshness
100/100How old is the benchmark?
Contamination history: AIME2024 heavily contaminated; 2025 cleaner but absorbs quickly (MathArena tracks fresh)
Limitations: Tiny (15-30) high-variance; one problem = ~3-7pts; guessable; 2024 contaminated, 2025 leaks fast
Who leads AIME
| Model | Score | Evidence |
|---|---|---|
| DeepSeek-R1 | 79.8 | self-reported |
| DeepSeek-R1-Distill-Qwen-32B | 72.6 | self-reported |
| DeepSeek-R1-Distill-Llama-70B | 70 | self-reported |
| DeepSeek-R1-Distill-Qwen-14B | 69.7 | self-reported |
| DeepSeek-R1-Distill-Qwen-7B | 55.5 | self-reported |
Frequently asked questions
AIME (AIME (2024/2025 as LLM benchmark)): AIME (2024/2025 as LLM benchmark) — HS competition math reasoning, integer answers; scored as Accuracy; often avg@k/cons@k.
DeepSeek-R1 leads AIME at 79.8 — self-reported by the lab. The full leaderboard above lists every recorded measurement, not just the headline number.
8 models have recorded scores on AIME, spanning a score spread of 50.9.
tensor.news grades AIME A for integrity (score 97/100), ranking #18 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — AIME still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is public, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Scores are reported under a consistent harness, so comparisons on AIME are reasonably apples-to-apples.
Every AIME measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites matharena.ai.