OpenAI·released 2024-09-121 source
o1-preview benchmark scores: 8 benchmarks tracked. 38% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 8 benchmarks·best result 81.65 on MATH level 5 (reproduced)·3 of 8 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MATH level 5 | 81.65accuracy (%) | reproduced· optimizedT1 | 2024-09-12Epoch AI |
| METR Time Horizons | 49.3accuracy (%) | unverifiedT1 | 2024-09-12METR - Measuring AI Ability to Complete Long Tasks |
| WeirdML | 47.56accuracy (%) | unverifiedT1 | 2024-09-12WeirdML Leaderboard |
| GPQA diamond | 33.75accuracy (%) | reproduced· optimizedT1 | 2024-09-12Epoch AI |
| OTIS Mock AIME 2024-2025 | 31.04accuracy (%) | reproduced· optimizedT1 | 2024-09-12Epoch AI |
| SimpleBench | 30.04accuracy (%) | unverifiedT1 | 2024-09-12SimpleBench Leaderboard |
| ARC-AGI | 18accuracy (%) | unverified· optimizedT1 | 2024-09-12ARC Prize Leaderboard |
| Cybench | 10accuracy (%) | unverified· optimizedT1 | 2024-09-12Cybench leaderboard |
Claims drawn from cited facts, not live model generation.
This model was developed by OpenAI and released in 2024.
2 cited facts
This model is scored on eight tracked benchmarks and currently holds no top score on any of them. Because the record has been independently reproduced, these scores can be treated as verified rather than merely vendor-claimed.
3 cited facts
Based on the available benchmark average, this model trails the leading score by about 53.81 points, a wide gap that places it well behind the front-of-pack. It also holds the top score on none of the tracked benchmarks, so there is no current category leadership.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
o1-preview is an AI model developed by OpenAI, released 2024-09-12. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
o1-preview has recorded scores on 8 benchmarks, each shown with its evidence status.
o1-preview has recorded scores on 8 benchmarks — GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, Cybench, SimpleBench, and 2 more. The full table above shows each score with its evidence status.
3 of 8 recorded scores (38%) are independently reproduced rather than self-reported by the lab.
o1-preview has 8 tracked claims: 3 independently reproduced, 5 unverified.