OpenAI·released 2026-03-171 source
GPT-5.4 Nano benchmark scores: 12 benchmarks tracked. 50% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 12 benchmarks·best result 87.77 on OTIS Mock AIME 2024-2025 (reproduced)·6 of 12 independently reproduced·$0.2/$1.25 per M tokens
Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 87.77accuracy (%) | reproduced· optimizedT1 | 2026-03-17Epoch AI |
| GPQA diamond | 71.29accuracy (%) | reproduced· optimizedT1 | 2026-03-17Epoch AI |
| ARC-AGI | 51.5accuracy (%) | unverified· optimizedT1 | 2026-03-17https://arcprize.org/leaderboard |
| WeirdML | 49.23accuracy (%) | unverifiedT1 | 2026-03-17https://htihle.github.io/weirdml.html |
| FrontierMath-Tiers-1-3-v2-Private | 44.91accuracy (%) | reproducedT1 | 2026-03-17Epoch AI |
| Chess Puzzles | 26.35accuracy (%) | reproducedT1 | 2026-03-17Epoch AI |
| APEX-Agents | 16.9accuracy (%) | unverifiedT2 | 2026-03-17 |
| FrontierMath-Tier-4-v2-Private | 12.2accuracy (%) | reproducedT1 | 2026-03-17Epoch AI |
| SimpleQA Verified | 12accuracy (%) | reproducedT1 | 2026-03-17Epoch AI |
| CritPt | 9.25accuracy (%) | unverifiedT2 | 2026-03-17 |
| ARC-AGI-2 | 5.69accuracy (%) | unverifiedT2 | 2026-03-17 |
| ProofBench | 5accuracy (%) | unverifiedT1 | 2026-03-17https://www.vals.ai/benchmarks/proof_bench |
Claims drawn from cited facts, not live model generation.
This model was developed by OpenAI and released in 2026.
2 cited facts
This model is scored on 12 tracked benchmarks, though it holds no current top score among them. Its benchmark record is independently reproduced, meaning outside evaluation backs the reported scores.
3 cited facts
With an average gap of 48.05 points behind the leader on the disclosed benchmark harness, this model sits in the wide-gap band — well behind the front-of-pack. It holds the top score on none of the surveyed benchmarks, so it has no verified category leadership.
2 cited facts
6 facts cross-checked across data sources: 2 corroborated, 2 single-source, 2 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, modelsdev_models, openrouter_models
GPT-5.4 Nano is an AI model developed by OpenAI, released 2026-03-17. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-5.4 Nano has recorded scores on 12 benchmarks, each shown with its evidence status.
GPT-5.4 Nano has recorded scores on 12 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleQA Verified, CritPt, and 6 more. The full table above shows each score with its evidence status.
6 of 12 recorded scores (50%) are independently reproduced rather than self-reported by the lab.
GPT-5.4 Nano has 12 tracked claims: 6 independently reproduced, 6 unverified.
Listed API pricing: $0.2 per million input tokens, $1.25 per million output tokens (prices disputed across sources). See the pricing block for the full breakdown.