OpenAI·released 2025-08-071 source
GPT-5 nano benchmark scores: 14 benchmarks tracked. 50% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 14 benchmarks·best result 95.24 on MATH level 5 (reproduced)·7 of 14 independently reproduced·$0.05/$0.4 per M tokens
Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MATH level 5 | 95.24accuracy (%) | reproduced· optimizedT1 | 2025-08-07Epoch AI |
| OTIS Mock AIME 2024-2025 | 81.09accuracy (%) | reproduced· optimizedT1 | 2025-08-07Epoch AI |
| GPQA diamond | 59.26accuracy (%) | reproduced· optimizedT1 | 2025-08-07Epoch AI |
| Fiction.LiveBench | 44.4accuracy (%) | unverifiedT1 | 2025-08-07Fiction.live leaderboard |
| WeirdML | 38.06accuracy (%) | unverifiedT1 | 2025-08-07WeirdML Leaderboard |
| Chess Puzzles | 23.19accuracy (%) | reproducedT1 | 2025-08-07Epoch AI |
| Terminal Bench | 21.8accuracy (%) | unverified· optimizedT1 | 2025-08-07https://www.tbench.ai/leaderboard/terminal-bench/2.0 |
| ARC-AGI | 20.71accuracy (%) | unverified· optimizedT2 | 2025-08-07 |
| FrontierMath-Tiers-1-3-v2-Private | 20accuracy (%) | reproducedT1 | 2025-08-07Epoch AI |
| SimpleQA Verified | 12.2accuracy (%) | reproducedT1 | 2025-08-07Epoch AI |
| ProofBench | 12accuracy (%) | unverifiedT1 | 2025-08-07https://www.vals.ai/benchmarks/proof_bench |
| VPCT | 5.8accuracy (%) | unverifiedT1 | 2025-08-07VPCT leaderboard |
| ARC-AGI-2 | 2.61accuracy (%) | unverifiedT2 | 2025-08-07 |
| FrontierMath-Tier-4-v2-Private | 2.44accuracy (%) | reproducedT1 | 2025-08-07Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by OpenAI and released in 2025.
2 cited facts
This model is evaluated on 14 tracked benchmarks, yet it currently holds no top score in any of them. Because the results are independently reproduced, the reported scores can be trusted as verified by outside evaluation.
3 cited facts
It trails the benchmark leader by an average of 58.51 points, a wide gap that places it well behind the front-runners rather than at the competitive frontier. It holds the top score on none of the tracked benchmarks, so it shows no category leadership yet.
2 cited facts
6 facts cross-checked across data sources: 2 corroborated, 2 single-source, 2 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, modelsdev_models, openrouter_models
GPT-5 nano is an AI model developed by OpenAI, released 2025-08-07. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-5 nano has recorded scores on 14 benchmarks, each shown with its evidence status.
GPT-5 nano has recorded scores on 14 benchmarks — GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleQA Verified, and 8 more. The full table above shows each score with its evidence status.
7 of 14 recorded scores (50%) are independently reproduced rather than self-reported by the lab.
GPT-5 nano has 14 tracked claims: 7 independently reproduced, 7 unverified.
Listed API pricing: $0.05 per million input tokens, $0.4 per million output tokens (prices disputed across sources). See the pricing block for the full breakdown.