OpenAI·released 2025-08-071 source
GPT-5 benchmark scores: 30 benchmarks tracked, leading 4. 33% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 4 of 30 benchmarks·best result 98.13 on MATH level 5 (reproduced)·10 of 30 independently reproduced·$1.25/$10 per M tokens
Head-to-headGPT-5 vs o3-proConsensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
Leads on: Aider polyglot, Fiction.LiveBench, Lech Mazur Writing, MATH level 5
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MATH level 5 | 98.13accuracy (%) | reproduced· optimizedT1 | 2025-08-07Epoch AI |
| Fiction.LiveBench | 97.2accuracy (%) | unverifiedT1 | 2025-08-07Fiction.live leaderboard |
| OTIS Mock AIME 2024-2025 | 91.38accuracy (%) | reproduced· optimizedT1 | 2025-08-07Epoch AI |
| Aider polyglot | 88accuracy (%) | unverified· optimizedT1 | 2025-08-07Aider LLM Leaderboards |
| Lech Mazur Writing | 86accuracy (%) | unverifiedT1 | 2025-08-07lechmazur/writing Github repository |
| GPQA diamond | 81.57accuracy (%) | reproduced· optimizedT1 | 2025-08-07Epoch AI |
| GeoBench | 81accuracy (%) | unverifiedT1 | 2025-08-07GeoBench leaderboard |
| SWE-Bench verified | 73.55accuracy (%) | reproduced· optimizedT1 | 2025-08-07Epoch AI |
| METR Time Horizons | 69.61accuracy (%) | unverifiedT1 | 2025-08-07METR - Measuring AI Ability to Complete Long Tasks |
| ARC-AGI | 65.7accuracy (%) | unverified· optimizedT1 | 2025-08-07ARC Prize Leaderboard |
| WeirdML | 60.7accuracy (%) | unverifiedT1 | 2025-08-07WeirdML Leaderboard |
| FrontierMath-Tiers-1-3-v2-Private | 55.44accuracy (%) | reproducedT1 | 2025-08-07Epoch AI |
| DeepResearch Bench | 55.13accuracy (%) | unverifiedT1 | 2025-08-07DeepResearchBench Leaderboard |
| SimpleQA Verified | 50.6accuracy (%) | reproducedT1 | 2025-08-07Epoch AI |
| Terminal Bench | 49.6accuracy (%) | unverified· optimizedT1 | 2025-08-07Terminal-Bench v2 Leaderboard |
| VPCT | 49accuracy (%) | unverifiedT1 | 2025-08-07VPCT leaderboard |
| SimpleBench | 48.04accuracy (%) | unverifiedT1 | 2025-08-07SimpleBench Leaderboard |
| GDPval | 34.8accuracy (%) | unverifiedT2 | 2025-08-07 |
| Chess Puzzles | 33.71accuracy (%) | reproducedT1 | 2025-08-07Epoch AI |
| Balrog | 32.8accuracy (%) | unverifiedT1 | 2025-08-07Balrog Leaderboard |
| FrontierMath-Tier-4-v2-Private | 21.95accuracy (%) | reproducedT1 | 2025-08-07Epoch AI |
| HLE | 21.55accuracy (%) | unverifiedT2 | 2025-08-07 |
| APEX-Agents | 18.3accuracy (%) | unverifiedT2 | 2025-08-07 |
| ProofBench | 18accuracy (%) | unverifiedT1 | 2025-08-07https://www.vals.ai/benchmarks/proof_bench |
| Mystery Game Puzzles | 15.18accuracy (%) | reproducedT1 | 2025-08-07Epoch AI |
| EBR-bench | 12.7accuracy (%) | reproducedT1 | 2025-08-07Epoch AI |
| CritPt | 12.6accuracy (%) | unverifiedT2 | 2025-08-07 |
| ARC-AGI-2 | 9.86accuracy (%) | unverifiedT2 | 2025-08-07 |
| GSO-Bench | 6.9accuracy (%) | unverifiedT1 | 2025-08-07GSO Leaderboard |
| Remote Labor Index | 1.67accuracy (%) | unverifiedT2 | 2025-08-07 |
Claims drawn from cited facts, not live model generation.
This model was developed by OpenAI and released in 2025.
2 cited facts
This model is tracked across 30 benchmarks and currently holds the top score on 4 of them. Because these results have been independently reproduced, the record is trustworthy.
3 cited facts
Across the disclosed harness, this model trails the average SOTA score by 25.73 points, a wide gap that places it well behind the leaders. It does top the leaderboard on 4 benchmarks, but that does not offset the large average deficit.
2 cited facts
6 facts cross-checked across data sources: 2 corroborated, 2 single-source, 2 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, modelsdev_models, openrouter_models
GPT-5 is an AI model developed by OpenAI, released 2025-08-07. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-5 has recorded scores on 30 benchmarks, each shown with its evidence status.
GPT-5 has recorded scores on 30 benchmarks — Lech Mazur Writing, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, and 24 more. The full table above shows each score with its evidence status.
10 of 30 recorded scores (33%) are independently reproduced rather than self-reported by the lab.
GPT-5 has 30 tracked claims: 10 independently reproduced, 20 unverified.
Listed API pricing: $1.25 per million input tokens, $10 per million output tokens (prices disputed across sources). See the pricing block for the full breakdown.