OpenAI·released 2026-03-051 source
GPT-5.4 Pro benchmark scores: 11 benchmarks tracked. 45% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 11 benchmarks·best result 94.5 on ARC-AGI (unverified)·5 of 11 independently reproduced·$30/$180 per M tokens
Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| ARC-AGI | 94.5accuracy (%) | unverified· optimizedT1 | 2026-03-05https://arcprize.org/leaderboard |
| GPQA diamond | 92.8accuracy (%) | reproduced· optimizedT1 | 2026-03-05Epoch AI |
| ARC-AGI-2 | 83.33accuracy (%) | unverifiedT2 | 2026-03-05 |
| FrontierMath-Tiers-1-3-v2-Private | 82.46accuracy (%) | reproducedT1 | 2026-03-05Epoch AI |
| SimpleBench | 68.92accuracy (%) | unverifiedT1 | 2026-03-05SimpleBench Leaderboard |
| FrontierMath-Tier-4-v2-Private | 58.54accuracy (%) | reproducedT1 | 2026-03-05Epoch AI |
| WeirdML | 57.44accuracy (%) | unverifiedT1 | 2026-03-05https://htihle.github.io/weirdml.html |
| Chess Puzzles | 56.44accuracy (%) | reproducedT1 | 2026-03-05Epoch AI |
| SimpleQA Verified | 47.8accuracy (%) | reproducedT1 | 2026-03-05Epoch AI |
| HLE | 41.51accuracy (%) | unverifiedT2 | 2026-03-05 |
| CritPt | 30accuracy (%) | unverifiedT2 | 2026-03-05 |
Claims drawn from cited facts, not live model generation.
This model was developed by OpenAI and released in 2026.
2 cited facts
This model is scored on 11 tracked benchmarks and currently holds no top score on any of them. Because the scores are independently reproduced, the record should be read as verified rather than vendor-claimed.
3 cited facts
On average, this model trails the benchmark leader by 12.09 points (under the disclosed harness), a wide gap that places it well behind the front-of-pack. It holds the top score on none of the tracked benchmarks, so it shows no genuine category leadership.
2 cited facts
6 facts cross-checked across data sources: 3 corroborated, 2 single-source, 1 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, modelsdev_models, openrouter_models
GPT-5.4 Pro is an AI model developed by OpenAI, released 2026-03-05. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-5.4 Pro has recorded scores on 11 benchmarks, each shown with its evidence status.
GPT-5.4 Pro has recorded scores on 11 benchmarks — GPQA diamond, WeirdML, Chess Puzzles, SimpleBench, SimpleQA Verified, CritPt, and 5 more. The full table above shows each score with its evidence status.
5 of 11 recorded scores (45%) are independently reproduced rather than self-reported by the lab.
GPT-5.4 Pro has 11 tracked claims: 5 independently reproduced, 6 unverified.
Listed API pricing: $30 per million input tokens, $180 per million output tokens. See the pricing block for the full breakdown.