OpenAI·released 2025-10-071 source
GPT-5 Pro benchmark scores: 7 benchmarks tracked. 29% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 7 benchmarks·best result 70.2 on ARC-AGI (unverified)·2 of 7 independently reproduced·$15/$120 per M tokens
Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| ARC-AGI | 70.2accuracy (%) | unverified· optimizedT1 | 2025-10-07ARC Prize Leaderboard |
| WeirdML | 60.39accuracy (%) | unverifiedT1 | 2025-10-07WeirdML Leaderboard |
| FrontierMath-Tiers-1-3-v2-Private | 55.79accuracy (%) | reproducedT1 | 2025-10-07Epoch AI |
| SimpleBench | 53.92accuracy (%) | unverifiedT1 | 2025-10-07SimpleBench Leaderboard |
| HLE | 28.19accuracy (%) | unverifiedT2 | 2025-10-07 |
| FrontierMath-Tier-4-v2-Private | 19.51accuracy (%) | reproducedT1 | 2025-10-07Epoch AI |
| ARC-AGI-2 | 18.33accuracy (%) | unverifiedT2 | 2025-10-07 |
Claims drawn from cited facts, not live model generation.
This model was developed by OpenAI and released in 2025.
2 cited facts
This model is evaluated on seven tracked benchmarks. It currently holds no top score on any of them. Because the results have been independently reproduced, the record is trustworthy.
3 cited facts
On average it trails the benchmark leader by 39.36 points, a wide gap that places it well behind the leaders. It holds the top score on none of the benchmarks, so it shows no category leadership yet.
2 cited facts
6 facts cross-checked across data sources: 3 corroborated, 2 single-source, 1 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, modelsdev_models, openrouter_models
GPT-5 Pro is an AI model developed by OpenAI, released 2025-10-07. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-5 Pro has recorded scores on 7 benchmarks, each shown with its evidence status.
GPT-5 Pro has recorded scores on 7 benchmarks — WeirdML, SimpleBench, ARC-AGI, ARC-AGI-2, FrontierMath-Tiers-1-3-v2-Private, HLE, and 1 more. The full table above shows each score with its evidence status.
2 of 7 recorded scores (29%) are independently reproduced rather than self-reported by the lab.
GPT-5 Pro has 7 tracked claims: 2 independently reproduced, 5 unverified.
Listed API pricing: $15 per million input tokens, $120 per million output tokens. See the pricing block for the full breakdown.