OpenAI·released 2025-08-071 source
GPT-5 mini benchmark scores: 18 benchmarks tracked. 44% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 18 benchmarks·best result 97.85 on MATH level 5 (reproduced)·8 of 18 independently reproduced·$0.25/$2 per M tokens
Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MATH level 5 | 97.85accuracy (%) | reproduced· optimizedT1 | 2025-08-07Epoch AI |
| OTIS Mock AIME 2024-2025 | 86.65accuracy (%) | reproduced· optimizedT1 | 2025-08-07Epoch AI |
| Lech Mazur Writing | 83.1accuracy (%) | unverifiedT1 | 2025-08-07lechmazur/writing Github repository |
| Fiction.LiveBench | 69.4accuracy (%) | unverifiedT1 | 2025-08-07Fiction.live leaderboard |
| GPQA diamond | 66.67accuracy (%) | reproduced· optimizedT1 | 2025-08-07Epoch AI |
| SWE-Bench verified | 64.67accuracy (%) | reproduced· optimizedT1 | 2025-08-07Epoch AI |
| ARC-AGI | 54.33accuracy (%) | unverified· optimizedT2 | 2025-08-07 |
| WeirdML | 52.67accuracy (%) | unverifiedT1 | 2025-08-07WeirdML Leaderboard |
| FrontierMath-Tiers-1-3-v2-Private | 46.67accuracy (%) | reproducedT1 | 2025-08-07Epoch AI |
| Terminal Bench | 34.8accuracy (%) | unverified· optimizedT1 | 2025-08-07https://www.tbench.ai/leaderboard/terminal-bench/2.0 |
| Chess Puzzles | 26.35accuracy (%) | reproducedT1 | 2025-08-07Epoch AI |
| SimpleQA Verified | 21accuracy (%) | reproducedT1 | 2025-08-07Epoch AI |
| HLE | 15.38accuracy (%) | unverifiedT2 | 2025-08-07 |
| FrontierMath-Tier-4-v2-Private | 12.2accuracy (%) | reproducedT1 | 2025-08-07Epoch AI |
| VPCT | 10.3accuracy (%) | unverifiedT1 | 2025-08-07VPCT leaderboard |
| ProofBench | 9accuracy (%) | unverifiedT1 | 2025-08-07https://www.vals.ai/benchmarks/proof_bench |
| ARC-AGI-2 | 4.44accuracy (%) | unverifiedT2 | 2025-08-07 |
| CritPt | 0accuracy (%) | unverifiedT2 | 2025-08-07 |
Claims drawn from cited facts, not live model generation.
Developed by OpenAI, this model was released in 2025.
2 cited facts
This model is tracked across 18 benchmarks. It currently holds no top score on any of them. The reported results are independently reproduced, so outside evaluation supports these scores.
3 cited facts
On the disclosed harness, this model trails the leader by an average of 41.55 points, a wide gap that places it well behind the front-of-pack. It holds the top score on none of the tracked benchmarks, so it has no measured category leadership yet.
2 cited facts
6 facts cross-checked across data sources: 2 corroborated, 2 single-source, 2 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, modelsdev_models, openrouter_models
GPT-5 mini is an AI model developed by OpenAI, released 2025-08-07. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-5 mini has recorded scores on 18 benchmarks, each shown with its evidence status.
GPT-5 mini has recorded scores on 18 benchmarks — Lech Mazur Writing, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, and 12 more. The full table above shows each score with its evidence status.
8 of 18 recorded scores (44%) are independently reproduced rather than self-reported by the lab.
GPT-5 mini has 18 tracked claims: 8 independently reproduced, 10 unverified.
Listed API pricing: $0.25 per million input tokens, $2 per million output tokens (prices disputed across sources). See the pricing block for the full breakdown.