OpenAI·released 2025-01-311 source
o3-mini benchmark scores: 17 benchmarks tracked. 35% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 17 benchmarks·best result 96.49 on MATH level 5 (reproduced)·6 of 17 independently reproduced·$1.1/$4.4 per M tokens
Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MATH level 5 | 96.49accuracy (%) | reproduced· optimizedT1 | 2025-01-31Epoch AI |
| OTIS Mock AIME 2024-2025 | 76.92accuracy (%) | reproduced· optimizedT1 | 2025-01-31Epoch AI |
| GPQA diamond | 69.36accuracy (%) | reproduced· optimizedT1 | 2025-01-31Epoch AI |
| Lech Mazur Writing | 61.7accuracy (%) | unverifiedT1 | 2025-01-31lechmazur/writing Github repository |
| Aider polyglot | 60.4accuracy (%) | unverified· optimizedT1 | 2025-01-31Aider LLM Leaderboards |
| CadEval | 54accuracy (%) | unverifiedT1 | 2025-01-31CadEval Dashboard |
| Fiction.LiveBench | 50accuracy (%) | unverifiedT1 | 2025-01-31Fiction.live leaderboard |
| WeirdML | 43.7accuracy (%) | unverifiedT1 | 2025-01-31WeirdML Leaderboard |
| ARC-AGI | 34.5accuracy (%) | unverified· optimizedT1 | 2025-01-31ARC Prize Leaderboard |
| Cybench | 22.5accuracy (%) | unverified· optimizedT1 | 2025-01-31Cybench leaderboard |
| FrontierMath-Tiers-1-3-v2-Private | 18.6accuracy (%) | reproducedT1 | 2025-01-31Epoch AI |
| Chess Puzzles | 12.67accuracy (%) | reproducedT1 | 2025-01-31Epoch AI |
| SimpleBench | 7.36accuracy (%) | unverifiedT1 | 2025-01-31SimpleBench Leaderboard |
| ARC-AGI-2 | 2.99accuracy (%) | unverifiedT2 | 2025-01-31 |
| GSO-Bench | 1.3accuracy (%) | unverifiedT1 | 2025-01-31GSO Leaderboard |
| CritPt | 0.29accuracy (%) | unverifiedT2 | 2025-01-31 |
| FrontierMath-Tier-4-v2-Private | 0accuracy (%) | reproducedT1 | 2025-01-31Epoch AI |
Claims drawn from cited facts, not live model generation.
OpenAI developed this model, and it was released in 2025. The developer has not disclosed a parameter count, so this model's scale cannot be characterized.
3 cited facts
This model is scored on 17 tracked benchmarks and currently holds no top score, as it leads none of them. Because these results have been independently reproduced, readers can treat the benchmark record as verified rather than merely vendor-claimed.
3 cited facts
This model trails the benchmark leader by an average of 46.84 points under the disclosed evaluation harness, a wide gap that places it well behind the front-of-pack. It holds the top score on none of the tracked benchmarks, meaning it has no category leadership yet.
2 cited facts
6 facts cross-checked across data sources: 3 corroborated, 1 single-source, 2 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, modelsdev_models, openrouter_models
o3-mini is an AI model developed by OpenAI, released 2025-01-31. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
o3-mini has recorded scores on 17 benchmarks, each shown with its evidence status.
o3-mini has recorded scores on 17 benchmarks — Lech Mazur Writing, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, CadEval, and 11 more. The full table above shows each score with its evidence status.
6 of 17 recorded scores (35%) are independently reproduced rather than self-reported by the lab.
o3-mini has 17 tracked claims: 6 independently reproduced, 11 unverified.
Listed API pricing: $1.1 per million input tokens, $4.4 per million output tokens. See the pricing block for the full breakdown.