OpenAI·released 2024-09-121 source
o1-mini benchmark scores: 11 benchmarks tracked. 36% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 11 benchmarks·best result 89.18 on MATH level 5 (reproduced)·4 of 11 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MATH level 5 | 89.18accuracy (%) | reproduced· optimizedT1 | 2024-09-12Epoch AI |
| Lech Mazur Writing | 64.9accuracy (%) | unverifiedT1 | 2024-09-12lechmazur/writing Github repository |
| GPQA diamond | 49.83accuracy (%) | reproduced· optimizedT1 | 2024-09-12Epoch AI |
| OTIS Mock AIME 2024-2025 | 46.89accuracy (%) | reproduced· optimizedT1 | 2024-09-12Epoch AI |
| WeirdML | 36.32accuracy (%) | unverifiedT1 | 2024-09-12WeirdML Leaderboard |
| Aider polyglot | 32.9accuracy (%) | unverified· optimizedT1 | 2024-09-12Aider LLM Leaderboards |
| ARC-AGI | 14accuracy (%) | unverified· optimizedT1 | 2024-09-12ARC Prize Leaderboard |
| Cybench | 10accuracy (%) | unverified· optimizedT1 | 2024-09-12Cybench leaderboard |
| FrontierMath-2025-02-28-Private | 3.02accuracy (%) | reproduced· optimizedT1 | 2024-09-12Epoch AI |
| SimpleBench | 1.72accuracy (%) | unverifiedT1 | 2024-09-12SimpleBench Leaderboard |
| ARC-AGI-2 | 0.83accuracy (%) | unverifiedT2 | 2024-09-12 |
Claims drawn from cited facts, not live model generation.
This model is developed by OpenAI and was released in 2024.
2 cited facts
Scored on 11 tracked benchmarks, this model holds no current top score, and because its record is independently reproduced, these scores are backed by outside evaluation and can be trusted.
3 cited facts
With an average gap of 57.02 points behind the state-of-the-art, this model trails well behind the leaders on its benchmarks. Binary bookkeeping is what a leader count amounts to — this model banks none of it, and distance to the front, the quantity that would matter, is beyond what the figure can express.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
o1-mini is an AI model developed by OpenAI, released 2024-09-12. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
o1-mini has recorded scores on 11 benchmarks, each shown with its evidence status.
o1-mini has recorded scores on 11 benchmarks — Lech Mazur Writing, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, Cybench, and 5 more. The full table above shows each score with its evidence status.
4 of 11 recorded scores (36%) are independently reproduced rather than self-reported by the lab.
o1-mini has 11 tracked claims: 4 independently reproduced, 7 unverified.