OpenAI·released 2025-06-101 source
o3-pro benchmark scores: 6 benchmarks tracked, leading 1. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 1 of 6 benchmarks·best result 97.2 on Fiction.LiveBench (unverified)·0 of 6 independently reproduced·$20/$80 per M tokens
Head-to-heado3-pro vs GPT-5Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
Leads on: Fiction.LiveBench
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| Fiction.LiveBench | 97.2accuracy (%) | unverifiedT1 | 2025-06-10Fiction.live leaderboard |
| Aider polyglot | 84.9accuracy (%) | unverified· optimizedT1 | 2025-06-10Aider LLM Leaderboards |
| Lech Mazur Writing | 84.4accuracy (%) | unverifiedT1 | 2025-06-10lechmazur/writing Github repository |
| ARC-AGI | 59.3accuracy (%) | unverified· optimizedT1 | 2025-06-10ARC Prize Leaderboard |
| WeirdML | 58.21accuracy (%) | unverifiedT1 | 2025-06-10WeirdML Leaderboard |
| ARC-AGI-2 | 4.86accuracy (%) | unverifiedT2 | 2025-06-10 |
Claims drawn from cited facts, not live model generation.
This model was developed by OpenAI and released in 2025.
2 cited facts
Topping 1 of 6 tracked benchmarks is a lead built on a small base — real, but too few measurements to generalize beyond those harnesses. Because its scores are unverified—neither independently reproduced nor formally contradicted—they should be treated as vendor-claimed results rather than confirmed facts.
3 cited facts
This model trails the state-of-the-art by an average of 26.93 points across benchmark evaluations, a wide gap that places it well behind the leaders. It holds the top score on one benchmark, indicating isolated strength but no overall frontier status.
2 cited facts
6 facts cross-checked across data sources: 3 corroborated, 2 single-source, 1 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, modelsdev_models, openrouter_models
o3-pro is an AI model developed by OpenAI, released 2025-06-10. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
o3-pro has recorded scores on 6 benchmarks, each shown with its evidence status.
o3-pro has recorded scores on 6 benchmarks — Lech Mazur Writing, WeirdML, Aider polyglot, ARC-AGI, Fiction.LiveBench, ARC-AGI-2. The full table above shows each score with its evidence status.
0 of 6 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
o3-pro has 6 tracked claims: 6 unverified.
Listed API pricing: $20 per million input tokens, $80 per million output tokens. See the pricing block for the full breakdown.