OpenAI·released 2025-04-161 source
o4-mini benchmark scores: 22 benchmarks tracked. 32% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 22 benchmarks·best result 97.83 on MATH level 5 (reproduced)·7 of 22 independently reproduced·$1.1/$4.4 per M tokens
Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MATH level 5 | 97.83accuracy (%) | reproduced· optimizedT1 | 2025-04-16Epoch AI |
| OTIS Mock AIME 2024-2025 | 81.65accuracy (%) | reproduced· optimizedT1 | 2025-04-16Epoch AI |
| Fiction.LiveBench | 77.8accuracy (%) | unverifiedT1 | 2025-04-16Fiction.live leaderboard |
| Lech Mazur Writing | 75accuracy (%) | unverifiedT1 | 2025-04-16lechmazur/writing Github repository |
| GPQA diamond | 72.81accuracy (%) | reproduced· optimizedT1 | 2025-04-16Epoch AI |
| Aider polyglot | 72accuracy (%) | unverified· optimizedT1 | 2025-04-16Aider LLM Leaderboards |
| GeoBench | 64accuracy (%) | unverifiedT1 | 2025-04-16GeoBench leaderboard |
| METR Time Horizons | 63.93accuracy (%) | unverifiedT1 | 2025-04-16METR - Measuring AI Ability to Complete Long Tasks |
| CadEval | 62accuracy (%) | unverifiedT1 | 2025-04-16CadEval Dashboard |
| ARC-AGI | 58.7accuracy (%) | unverified· optimizedT1 | 2025-04-16ARC Prize Leaderboard |
| WeirdML | 52.56accuracy (%) | unverifiedT1 | 2025-04-16WeirdML Leaderboard |
| VPCT | 36.25accuracy (%) | unverifiedT1 | 2025-04-16VPCT leaderboard |
| FrontierMath-Tiers-1-3-v2-Private | 36.14accuracy (%) | reproducedT1 | 2025-04-16Epoch AI |
| SimpleBench | 26.44accuracy (%) | unverifiedT1 | 2025-04-16SimpleBench Leaderboard |
| GDPval | 25.3accuracy (%) | unverifiedT2 | 2025-04-16 |
| SimpleQA Verified | 23.9accuracy (%) | reproducedT1 | 2025-04-16Epoch AI |
| Chess Puzzles | 22.14accuracy (%) | reproducedT1 | 2025-04-16Epoch AI |
| HLE | 13.95accuracy (%) | unverifiedT2 | 2025-04-16 |
| ARC-AGI-2 | 6.11accuracy (%) | unverifiedT2 | 2025-04-16 |
| FrontierMath-Tier-4-v2-Private | 4.88accuracy (%) | reproducedT1 | 2025-04-16Epoch AI |
| GSO-Bench | 3.6accuracy (%) | unverifiedT1 | 2025-04-16GSO Leaderboard |
| CritPt | 0.6accuracy (%) | unverifiedT2 | 2025-04-16 |
Claims drawn from cited facts, not live model generation.
This model was developed by OpenAI and released in 2025.
2 cited facts
This model is tracked on 22 benchmarks and holds no current top score on any of them. Because its results have been independently reproduced, the record is externally corroborated rather than merely vendor-claimed.
3 cited facts
On average, this model trails the state-of-the-art leader by 34.66 points under the disclosed harness, a wide gap that places it well behind the front-runners. It additionally holds the top score on none of the tracked benchmarks, indicating no category leadership yet.
2 cited facts
6 facts cross-checked across data sources: 3 corroborated, 1 single-source, 2 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, modelsdev_models, openrouter_models
o4-mini is an AI model developed by OpenAI, released 2025-04-16. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
o4-mini has recorded scores on 22 benchmarks, each shown with its evidence status.
o4-mini has recorded scores on 22 benchmarks — Lech Mazur Writing, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, CadEval, and 16 more. The full table above shows each score with its evidence status.
7 of 22 recorded scores (32%) are independently reproduced rather than self-reported by the lab.
o4-mini has 22 tracked claims: 7 independently reproduced, 15 unverified.
Listed API pricing: $1.1 per million input tokens, $4.4 per million output tokens. See the pricing block for the full breakdown.