OpenAI·released 2026-03-051 source
GPT-5.4 benchmark scores: 25 benchmarks tracked, leading 1. 40% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 1 of 25 benchmarks·best result 97.78 on OTIS Mock AIME 2024-2025 (reproduced)·10 of 25 independently reproduced·$2.5/$15 per M tokens
Head-to-headGPT-5.4 vs GPT-5.1Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
Leads on: CL-bench
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 97.78accuracy (%) | reproduced· optimizedT1 | 2026-03-05Epoch AI |
| ARC-AGI | 93.67accuracy (%) | unverified· optimizedT1 | 2026-03-05https://arcprize.org/leaderboard |
| GPQA diamond | 91.07accuracy (%) | reproduced· optimizedT1 | 2026-03-05Epoch AI |
| Terminal Bench | 81.8accuracy (%) | unverified· optimizedT1 | 2026-03-05https://www.tbench.ai/leaderboard/terminal-bench/2.0 |
| FrontierMath-Tiers-1-3-v2-Private | 78.6accuracy (%) | reproducedT1 | 2026-03-05Epoch AI |
| WeirdML | 77.7accuracy (%) | unverifiedT1 | 2026-03-05WeirdML Leaderboard |
| SWE-Bench verified | 76.86accuracy (%) | reproduced· optimizedT1 | 2026-03-05Epoch AI |
| METR Time Horizons | 74.34accuracy (%) | self-reportedT1 | 2026-03-05https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ |
| ARC-AGI-2 | 73.95accuracy (%) | unverifiedT2 | 2026-03-05 |
| ProofBench | 56accuracy (%) | unverifiedT1 | 2026-03-05https://www.vals.ai/benchmarks/proof_bench |
| DeepSWE | 51.77accuracy (%) | unverifiedT1 | 2026-03-05https://deepswe.datacurve.ai/ |
| FrontierMath-Tier-4-v2-Private | 49accuracy (%) | reproducedT1 | 2026-03-05Epoch AI |
| SimpleQA Verified | 44.83accuracy (%) | reproducedT1 | 2026-03-05Epoch AI |
| Chess Puzzles | 41.08accuracy (%) | reproducedT1 | 2026-03-05Epoch AI |
| APEX-Agents | 36accuracy (%) | unverifiedT2 | 2026-03-05 |
| DeepResearch Bench | 35.09accuracy (%) | unverifiedT1 | 2026-03-05https://drb.futuresearch.ai/#drb |
| HLE | 33.03accuracy (%) | unverifiedT2 | 2026-03-05 |
| GSO-Bench | 31.37accuracy (%) | unverifiedT1 | 2026-03-05https://gso-bench.github.io/index.html |
| Mystery Game Puzzles | 30.6accuracy (%) | reproducedT1 | 2026-03-05Epoch AI |
| CL-bench | 27.9accuracy (%) | unverifiedT2 | 2026-03-05 |
| EBR-bench | 25.4accuracy (%) | reproducedT1 | 2026-03-05Epoch AI |
| CritPt | 23.43accuracy (%) | unverifiedT2 | 2026-03-05 |
| CL-bench Life | 21.7accuracy (%) | unverifiedT2 | 2026-03-05 |
| PostTrainBench | 20.23accuracy (%) | unverifiedT2 | 2026-03-05 |
| MirrorCode | 15.56accuracy (%) | reproducedT1 | 2026-03-05Epoch AI |
Claims drawn from cited facts, not live model generation.
This model is scored on 25 tracked benchmarks, and it currently holds the top score on 1 of them. Because these scores have been independently reproduced, they can be treated as verified results rather than vendor claims.
3 cited facts
On average, this model trails the state-of-the-art leader by 16.29 points under the disclosed harness, a wide gap that places it well behind the front of the pack. It does hold the top score on 1 benchmark, but that single leading result does not offset the average deficit, so genuine category leadership is not established overall.
2 cited facts
6 facts cross-checked across data sources: 3 corroborated, 2 single-source, 1 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, modelsdev_models, openrouter_models
GPT-5.4 is an AI model developed by OpenAI, released 2026-03-05. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-5.4 has recorded scores on 25 benchmarks, each shown with its evidence status.
GPT-5.4 has recorded scores on 25 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, METR Time Horizons, SimpleQA Verified, and 19 more. The full table above shows each score with its evidence status.
10 of 25 recorded scores (40%) are independently reproduced rather than self-reported by the lab.
GPT-5.4 has 25 tracked claims: 10 independently reproduced, 1 self-reported, 14 unverified.
Listed API pricing: $2.5 per million input tokens, $15 per million output tokens. See the pricing block for the full breakdown.