OpenAI·released 2025-11-131 source
GPT-5.1 benchmark scores: 20 benchmarks tracked. 35% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 20 benchmarks·best result 88.6 on OTIS Mock AIME 2024-2025 (reproduced)·7 of 20 independently reproduced·$1.25/$10 per M tokens
Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 88.6accuracy (%) | reproduced· optimizedT1 | 2025-11-13Epoch AI |
| GPQA diamond | 83.5accuracy (%) | reproduced· optimizedT1 | 2025-11-13Epoch AI |
| ARC-AGI | 72.83accuracy (%) | unverified· optimizedT1 | 2025-11-13https://arcprize.org/leaderboard |
| SWE-Bench verified | 67.98accuracy (%) | reproduced· optimizedT1 | 2025-11-13Epoch AI |
| WeirdML | 60.77accuracy (%) | unverifiedT1 | 2025-11-13WeirdML Leaderboard |
| FrontierMath-2025-02-28-Private | 54.45accuracy (%) | reproduced· optimizedT1 | 2025-11-13Epoch AI |
| SimpleQA Verified | 48.9accuracy (%) | reproducedT1 | 2025-11-13Epoch AI |
| Terminal Bench | 47.6accuracy (%) | unverified· optimizedT1 | 2025-11-13Terminal-Bench v2 Leaderboard |
| SimpleBench | 43.84accuracy (%) | unverifiedT1 | 2025-11-13SimpleBench Leaderboard |
| DeepResearch Bench | 42.79accuracy (%) | unverifiedT1 | 2025-11-13https://drb.futuresearch.ai/#drb |
| VPCT | 38.05accuracy (%) | unverifiedT1 | 2025-11-13VPCT leaderboard |
| Chess Puzzles | 28.45accuracy (%) | reproducedT1 | 2025-11-13Epoch AI |
| CL-bench | 23.7accuracy (%) | unverifiedT2 | 2025-11-13 |
| FrontierMath-Tier-4-2025-07-01-Private | 20.83accuracy (%) | reproducedT1 | 2025-11-13Epoch AI |
| HLE | 19.83accuracy (%) | unverifiedT2 | 2025-11-13 |
| ARC-AGI-2 | 17.64accuracy (%) | unverifiedT2 | 2025-11-13 |
| APEX-Agents | 17.5accuracy (%) | unverifiedT2 | 2025-11-13 |
| CL-bench Life | 17.3accuracy (%) | unverifiedT2 | 2025-11-13 |
| GSO-Bench | 13.73accuracy (%) | unverifiedT1 | 2025-11-13https://gso-bench.github.io/index.html |
| CritPt | 4.86accuracy (%) | unverifiedT2 | 2025-11-13 |
Claims drawn from cited facts, not live model generation.
This model originates from OpenAI and was released in 2025.
2 cited facts
This model is evaluated across 20 tracked benchmarks and currently holds no top score on any of them. Because the available scores have been independently reproduced, the record can be trusted as verified rather than treated as self-reported claims.
3 cited facts
Across the benchmark suite, the model trails the aggregate leader by an average of 25.42 points, a wide gap that places it well behind the front of the pack rather than merely competitive. It holds the top score on none of the benchmarks, meaning it has no genuine category leadership yet.
2 cited facts
6 facts cross-checked across data sources: 2 corroborated, 2 single-source, 2 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, modelsdev_models, openrouter_models
GPT-5.1 is an AI model developed by OpenAI, released 2025-11-13. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-5.1 has recorded scores on 20 benchmarks, each shown with its evidence status.
GPT-5.1 has recorded scores on 20 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleBench, FrontierMath-2025-02-28-Private, and 14 more. The full table above shows each score with its evidence status.
7 of 20 recorded scores (35%) are independently reproduced rather than self-reported by the lab.
GPT-5.1 has 20 tracked claims: 7 independently reproduced, 13 unverified.
Listed API pricing: $1.25 per million input tokens, $10 per million output tokens (prices disputed across sources). See the pricing block for the full breakdown.