Benchmark explainer
What is OTIS Mock AIME 2024-2025?
Competition-math reasoning on novel (un-leaked) AIME-style integer-answer problems crafted by elite olympiad students.
OTIS Mock AIME 2024-2025 as reported in Epoch AI's Capabilities Index CSV.
How it's scored
- Metric
- normalized performance in Epoch CSV, rendered as accuracy (%)
- Score ceiling
- 100
- Construction
- 45 problems across 3 exams: 15 from Mock AIME 2024 + 15 each from Mock AIME 2025 I & II; 3-hour exam, integer answers 0-999.
- Human baseline
- AIME qualifier threshold ~7-10/15 correct (~50%); Mock AIMEs typically harder by 2-4 problems (per Evan Chen).
How trustworthy
Discrimination
100/100Does it still separate models?
Saturation headroom
91.1/100How far from ceiling / clustered at the top?
Contamination resistance
26.6/100Public vs held-out; training-leak risk.
Harness comparability
100/100Is it apples-to-apples, or mixed / vendor-optimized?
Freshness
84.9/100How old is the benchmark?
Contamination history: Problems novel and unreleased before 2024-2025 — explicit Epoch rationale: less leakage risk than official AIME.
Limitations: Difficulty ranges from AIME-level to IMO-style (mixed artistic license); image-based problems excluded for parity; small set (~45) limits statistical power.
Who leads OTIS Mock AIME 2024-2025
| Model | Score | Evidence |
|---|---|---|
| Claude Fable 5 | 100 | reproduced |
| Claude Fable 5.1 | 100 | unverified |
| Qwen3.8 Max (0902) | 100 | unverified |
| GPT-5.5 | 100 | reproduced |
| GPT-5.5 Pro | 100 | reproduced |
Compare the top OTIS Mock AIME 2024-2025 scorers
Frequently asked questions
OTIS Mock AIME 2024-2025: OTIS Mock AIME 2024-2025 as reported in Epoch AI's Capabilities Index CSV.
Claude Fable 5 leads OTIS Mock AIME 2024-2025 at 100 — independently reproduced. The full leaderboard above lists every recorded measurement, not just the headline number.
169 models have recorded scores on OTIS Mock AIME 2024-2025, spanning a score spread of 100.
tensor.news grades OTIS Mock AIME 2024-2025 A for integrity (score 85/100), ranking #34 of 62 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
Yes — top scores are clustering near the ceiling (saturation ratio 0.09), so OTIS Mock AIME 2024-2025 no longer separates leading models well. Treat small gaps at the top with caution.
Its test set is unknown, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Scores are reported under a consistent harness, so comparisons on OTIS Mock AIME 2024-2025 are reasonably apples-to-apples.
Every OTIS Mock AIME 2024-2025 measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites epoch.ai/benchmarks/otis-mock-aime-2024-2025.