OpenAI·released 2024-12-201 source
o3 benchmark scores: 27 benchmarks tracked, leading 1. 30% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 1 of 27 benchmarks·best result 97.77 on MATH level 5 (reproduced)·8 of 27 independently reproduced·$2/$8 per M tokens
Head-to-heado3 vs Gemini 2.5 Pro (Mar 2025)Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
Leads on: CadEval
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MATH level 5 | 97.77accuracy (%) | reproduced· optimizedT1 | 2024-12-20Epoch AI |
| Fiction.LiveBench | 88.9accuracy (%) | unverifiedT1 | 2024-12-20Fiction.live leaderboard |
| OTIS Mock AIME 2024-2025 | 84.43accuracy (%) | reproduced· optimizedT1 | 2024-12-20Epoch AI |
| Lech Mazur Writing | 83.9accuracy (%) | unverifiedT1 | 2024-12-20lechmazur/writing Github repository |
| Aider polyglot | 81.3accuracy (%) | unverified· optimizedT1 | 2024-12-20Aider LLM Leaderboards |
| GPQA diamond | 75.76accuracy (%) | reproduced· optimizedT1 | 2024-12-20Epoch AI |
| CadEval | 74accuracy (%) | unverifiedT1 | 2024-12-20CadEval Dashboard |
| GeoBench | 74accuracy (%) | unverifiedT1 | 2024-12-20GeoBench leaderboard |
| METR Time Horizons | 65.44accuracy (%) | unverifiedT1 | 2024-12-20METR - Measuring AI Ability to Complete Long Tasks |
| SWE-Bench verified | 62.32accuracy (%) | reproduced· optimizedT1 | 2024-12-20Epoch AI |
| ARC-AGI | 60.8accuracy (%) | unverified· optimizedT1 | 2024-12-20ARC Prize Leaderboard |
| SimpleQA Verified | 53accuracy (%) | reproducedT1 | 2024-12-20Epoch AI |
| WeirdML | 52.42accuracy (%) | unverifiedT1 | 2024-12-20WeirdML Leaderboard |
| DeepResearch Bench | 46.6accuracy (%) | unverifiedT1 | 2024-12-20DeepResearchBench Leaderboard |
| SimpleBench | 43.72accuracy (%) | unverifiedT1 | 2024-12-20SimpleBench Leaderboard |
| Chess Puzzles | 34.76accuracy (%) | reproducedT1 | 2024-12-20Epoch AI |
| FrontierMath-2025-02-28-Private | 32.78accuracy (%) | reproduced· optimizedT1 | 2024-12-20Epoch AI |
| GDPval | 30.8accuracy (%) | unverifiedT2 | 2024-12-20 |
| VPCT | 28accuracy (%) | unverifiedT1 | 2024-12-20VPCT leaderboard |
| OSWorld | 23accuracy (%) | unverified· optimizedT1 | 2024-12-20OS World Website |
| CL-bench | 17.8accuracy (%) | unverifiedT2 | 2024-12-20 |
| APEX-Agents | 17.2accuracy (%) | unverifiedT2 | 2024-12-20 |
| HLE | 16.3accuracy (%) | unverifiedT2 | 2024-12-20 |
| GSO-Bench | 8.8accuracy (%) | unverifiedT1 | 2024-12-20GSO Leaderboard |
| ARC-AGI-2 | 6.53accuracy (%) | unverifiedT2 | 2024-12-20 |
| FrontierMath-Tier-4-2025-07-01-Private | 3.47accuracy (%) | reproducedT1 | 2024-12-20Epoch AI |
| CritPt | 1.4accuracy (%) | unverifiedT2 | 2024-12-20 |
Claims drawn from cited facts, not live model generation.
This model was developed by OpenAI and released in 2024.
2 cited facts
This model is scored on 27 tracked benchmarks and currently holds the top score on 1 of them. Because its scores have been independently reproduced, outside evaluation backs the record, so the results can be treated as verified rather than vendor claims.
3 cited facts
It trails the SOTA leader by an average of 25.24 points on its benchmarks, a wide gap that places it well behind the front of the pack. It holds the top score on 1 benchmark, a narrow instance of category leadership that does not offset the overall deficit.
2 cited facts
6 facts cross-checked across data sources: 3 corroborated, 2 single-source, 1 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, modelsdev_models, openrouter_models
o3 is an AI model developed by OpenAI, released 2024-12-20. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
o3 has recorded scores on 27 benchmarks, each shown with its evidence status.
o3 has recorded scores on 27 benchmarks — Lech Mazur Writing, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, CadEval, and 21 more. The full table above shows each score with its evidence status.
8 of 27 recorded scores (30%) are independently reproduced rather than self-reported by the lab.
o3 has 27 tracked claims: 8 independently reproduced, 19 unverified.
Listed API pricing: $2 per million input tokens, $8 per million output tokens. See the pricing block for the full breakdown.