OpenAI·released 2026-04-231 source
GPT-5.5 benchmark scores: 28 benchmarks tracked, leading 4. 36% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 4 of 28 benchmarks·best result 100 on OTIS Mock AIME 2024-2025 (reproduced)·10 of 28 independently reproduced·$5/$30 per M tokens
Head-to-headGPT-5.5 vs GPT-5.4Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
Leads on: CL-bench Life, ExploitBench, OTIS Mock AIME 2024-2025, Terminal Bench
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 100accuracy (%) | reproduced· optimizedT1 | 2026-04-23Epoch AI |
| ARC-AGI | 95accuracy (%) | unverified· optimizedT1 | 2026-04-23https://arcprize.org/leaderboard |
| GPQA diamond | 92accuracy (%) | reproduced· optimizedT1 | 2026-04-23Epoch AI |
| Surface Evolver Bench | 88.12accuracy (%) | unverifiedT1 | 2026-04-23https://yhenon.github.io/surface-evolver-llm-eval/ |
| FrontierMath-Tiers-1-3-v2-Private | 85.26accuracy (%) | reproducedT1 | 2026-04-23Epoch AI |
| ARC-AGI-2 | 85accuracy (%) | unverifiedT2 | 2026-04-23 |
| WeirdML | 84.91accuracy (%) | unverifiedT1 | 2026-04-23https://htihle.github.io/weirdml.html |
| Terminal Bench | 84.7accuracy (%) | unverified· optimizedT1 | 2026-04-23https://www.tbench.ai/leaderboard/terminal-bench/2.0 |
| SWE-Bench verified | 80.58accuracy (%) | reproduced· optimizedT1 | 2026-04-23Epoch AI |
| FrontierMath-Tier-4-v2-Private | 72.5accuracy (%) | reproducedT1 | 2026-04-23Epoch AI |
| DeepSWE | 67.04accuracy (%) | unverifiedT1 | 2026-04-23https://deepswe.datacurve.ai/ |
| SimpleQA Verified | 63.1accuracy (%) | reproducedT1 | 2026-04-23Epoch AI |
| SimpleBench | 62.8accuracy (%) | unverifiedT1 | 2026-04-23SimpleBench Leaderboard |
| DeepResearch Bench | 54.01accuracy (%) | unverifiedT1 | 2026-04-23https://drb.futuresearch.ai/#drb |
| Chess Puzzles | 51.6accuracy (%) | reproducedT1 | 2026-04-23Epoch AI |
| Mystery Game Puzzles | 51.53accuracy (%) | reproducedT1 | 2026-04-23Epoch AI |
| ProofBench | 50accuracy (%) | unverifiedT1 | 2026-04-23https://www.vals.ai/benchmarks/proof_bench |
| ExploitBench | 47.4accuracy (%) | unverifiedT2 | 2026-04-23 |
| FrontierCode | 43accuracy (%) | unverifiedT1 | 2026-04-23https://cognition.com/frontiercode |
| GSO-Bench | 40.2accuracy (%) | unverifiedT1 | 2026-04-23GSO Leaderboard |
| APEX-Agents | 38.5accuracy (%) | unverifiedT2 | 2026-04-23 |
| EBR-bench | 34.29accuracy (%) | reproducedT1 | 2026-04-23Epoch AI |
| CritPt | 27.14accuracy (%) | unverifiedT2 | 2026-04-23 |
| PostTrainBench | 25.02accuracy (%) | unverifiedT2 | 2026-04-23 |
| CL-bench Life | 22.2accuracy (%) | unverifiedT2 | 2026-04-23 |
| OSWorld 2.0 | 13accuracy (%) | unverifiedT1 | 2026-04-23https://osworld-v2.xlang.ai/ |
| MirrorCode | 10accuracy (%) | reproducedT1 | 2026-04-23Epoch AI |
| Remote Labor Index | 6.25accuracy (%) | unverifiedT2 | 2026-04-23 |
Claims drawn from cited facts, not live model generation.
Developed by OpenAI, this model was released in 2026.
2 cited facts
This model has been evaluated across 28 tracked benchmarks and currently holds the top score on 4 of them. Because these results have been independently reproduced by outside evaluation, the scores can be treated as verified rather than merely vendor-claimed.
3 cited facts
It trails the SOTA leader by an average of 10.05 points under the disclosed harness, a wide gap that places it well behind the front of the pack. It still holds the top score on 4 benchmarks, so it can lead on specific tasks even while the average gap is large.
2 cited facts
6 facts cross-checked across data sources: 3 corroborated, 2 single-source, 1 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: curated_capability_claims, epoch_benchmarks, litellm_prices, modelsdev_models +1 more
GPT-5.5 is an AI model developed by OpenAI, released 2026-04-23. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-5.5 has recorded scores on 28 benchmarks, each shown with its evidence status.
GPT-5.5 has recorded scores on 28 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleBench, SimpleQA Verified, and 22 more. The full table above shows each score with its evidence status.
10 of 28 recorded scores (36%) are independently reproduced rather than self-reported by the lab.
GPT-5.5 has 28 tracked claims: 10 independently reproduced, 18 unverified.
Listed API pricing: $5 per million input tokens, $30 per million output tokens. See the pricing block for the full breakdown.