OpenAI·released 2024-12-051 source
o1 benchmark scores: 17 benchmarks tracked. 29% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 17 benchmarks·best result 94.71 on MATH level 5 (reproduced)·5 of 17 independently reproduced·$15/$60 per M tokens
Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MATH level 5 | 94.71accuracy (%) | reproduced· optimizedT1 | 2024-12-05Epoch AI |
| Fiction.LiveBench | 83.3accuracy (%) | unverifiedT1 | 2024-12-05Fiction.live leaderboard |
| GeoBench | 80accuracy (%) | unverifiedT1 | 2024-12-05GeoBench leaderboard |
| OTIS Mock AIME 2024-2025 | 73.31accuracy (%) | reproduced· optimizedT1 | 2024-12-05Epoch AI |
| Lech Mazur Writing | 70.2accuracy (%) | unverifiedT1 | 2024-12-05lechmazur/writing Github repository |
| GPQA diamond | 69.02accuracy (%) | reproduced· optimizedT1 | 2024-12-05Epoch AI |
| Aider polyglot | 61.7accuracy (%) | unverified· optimizedT1 | 2024-12-05Aider LLM Leaderboards |
| CadEval | 56accuracy (%) | unverifiedT1 | 2024-12-05CadEval Dashboard |
| METR Time Horizons | 55.93accuracy (%) | unverifiedT1 | 2024-12-05METR - Measuring AI Ability to Complete Long Tasks |
| WeirdML | 43.82accuracy (%) | unverifiedT1 | 2024-12-05WeirdML Leaderboard |
| ARC-AGI | 30.7accuracy (%) | unverified· optimizedT1 | 2024-12-05ARC Prize Leaderboard |
| SimpleBench | 28.12accuracy (%) | unverifiedT1 | 2024-12-05SimpleBench Leaderboard |
| FrontierMath-2025-02-28-Private | 16.33accuracy (%) | reproduced· optimizedT1 | 2024-12-05Epoch AI |
| Chess Puzzles | 10.56accuracy (%) | reproducedT1 | 2024-12-05Epoch AI |
| VPCT | 5.5accuracy (%) | unverifiedT1 | 2024-12-05VPCT leaderboard |
| HLE | 3.32accuracy (%) | unverifiedT2 | 2024-12-05 |
| APEX-Agents | 1.1accuracy (%) | unverifiedT2 | 2024-12-05 |
Claims drawn from cited facts, not live model generation.
This model was developed by OpenAI. It was released in 2024.
2 cited facts
This model is scored on 17 tracked benchmarks. It currently holds no top score among them. Its results are independently reproduced, so the record can be treated as trustworthy.
3 cited facts
With an average gap of 34.95 points behind the leading score under the disclosed harness, this model sits well behind the front of the pack. It holds the top score on none of the tracked benchmarks, so it currently shows no genuine category leadership.
2 cited facts
6 facts cross-checked across data sources: 3 corroborated, 1 single-source, 2 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, modelsdev_models, openrouter_models
o1 is an AI model developed by OpenAI, released 2024-12-05. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
o1 has recorded scores on 17 benchmarks, each shown with its evidence status.
o1 has recorded scores on 17 benchmarks — Lech Mazur Writing, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, CadEval, and 11 more. The full table above shows each score with its evidence status.
5 of 17 recorded scores (29%) are independently reproduced rather than self-reported by the lab.
o1 has 17 tracked claims: 5 independently reproduced, 12 unverified.
Listed API pricing: $15 per million input tokens, $60 per million output tokens. See the pricing block for the full breakdown.