OpenAI·released 2026-03-171 source
GPT-5.4 Mini benchmark scores: 14 benchmarks tracked. 43% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 14 benchmarks·best result 88.88 on OTIS Mock AIME 2024-2025 (reproduced)·6 of 14 independently reproduced·$0.75/$4.5 per M tokens
Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 88.88accuracy (%) | reproduced· optimizedT1 | 2026-03-17Epoch AI |
| GPQA diamond | 82.49accuracy (%) | reproduced· optimizedT1 | 2026-03-17Epoch AI |
| ARC-AGI | 63.67accuracy (%) | unverified· optimizedT1 | 2026-03-17https://arcprize.org/leaderboard |
| WeirdML | 60.3accuracy (%) | unverifiedT1 | 2026-03-17https://htihle.github.io/weirdml.html |
| FrontierMath-Tiers-1-3-v2-Private | 51.23accuracy (%) | reproducedT1 | 2026-03-17Epoch AI |
| DeepResearch Bench | 36.26accuracy (%) | unverifiedT1 | 2026-03-17https://drb.futuresearch.ai/#drb |
| SimpleQA Verified | 28.6accuracy (%) | reproducedT1 | 2026-03-17Epoch AI |
| FrontierCode | 27.04accuracy (%) | unverifiedT1 | 2026-03-17https://cognition.com/frontiercode |
| APEX-Agents | 24.6accuracy (%) | unverifiedT2 | 2026-03-17 |
| ProofBench | 21accuracy (%) | unverifiedT1 | 2026-03-17https://www.vals.ai/benchmarks/proof_bench |
| Chess Puzzles | 20.03accuracy (%) | reproducedT1 | 2026-03-17Epoch AI |
| ARC-AGI-2 | 18.9accuracy (%) | unverifiedT2 | 2026-03-17 |
| CritPt | 10accuracy (%) | unverifiedT2 | 2026-03-17 |
| FrontierMath-Tier-4-v2-Private | 9.76accuracy (%) | reproducedT1 | 2026-03-17Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by OpenAI. It was released in 2026, and with no parameter count available, its scale cannot be characterized.
2 cited facts
This model is tracked across fourteen benchmarks, though it currently holds no top score in any of them. Because its results have been independently reproduced, those scores can be treated as verified rather than vendor claims.
3 cited facts
On average across the disclosed benchmarks it trails the leader by 38.2 points, a wide gap that places it well behind the front of the pack. It holds the top score on none of the tracked benchmarks, so no genuine category leadership is currently indicated.
2 cited facts
6 facts cross-checked across data sources: 2 corroborated, 2 single-source, 2 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, modelsdev_models, openrouter_models
GPT-5.4 Mini is an AI model developed by OpenAI, released 2026-03-17. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-5.4 Mini has recorded scores on 14 benchmarks, each shown with its evidence status.
GPT-5.4 Mini has recorded scores on 14 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleQA Verified, CritPt, and 8 more. The full table above shows each score with its evidence status.
6 of 14 recorded scores (43%) are independently reproduced rather than self-reported by the lab.
GPT-5.4 Mini has 14 tracked claims: 6 independently reproduced, 8 unverified.
Listed API pricing: $0.75 per million input tokens, $4.5 per million output tokens (prices disputed across sources). See the pricing block for the full breakdown.