OpenAI·released 2026-04-231 source
GPT-5.5 Pro benchmark scores: 10 benchmarks tracked, leading 2. 60% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 2 of 10 benchmarks·best result 100 on OTIS Mock AIME 2024-2025 (reproduced)·6 of 10 independently reproduced·$30/$180 per M tokens
Head-to-headGPT-5.5 Pro vs GPT-5.6 SolConsensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
Leads on: Chess Puzzles, OTIS Mock AIME 2024-2025
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 100accuracy (%) | reproduced· optimizedT1 | 2026-04-23Epoch AI |
| ARC-AGI | 96.5accuracy (%) | unverified· optimizedT1 | 2026-04-23https://arcprize.org/leaderboard |
| GPQA diamond | 91.9accuracy (%) | reproduced· optimizedT1 | 2026-04-23Epoch AI |
| FrontierMath-Tiers-1-3-v2-Private | 87.72accuracy (%) | reproducedT1 | 2026-04-23Epoch AI |
| ARC-AGI-2 | 84.58accuracy (%) | unverifiedT2 | 2026-04-23 |
| FrontierMath-Tier-4-v2-Private | 78.05accuracy (%) | reproducedT1 | 2026-04-23Epoch AI |
| SimpleBench | 72.28accuracy (%) | unverifiedT1 | 2026-04-23SimpleBench Leaderboard |
| SimpleQA Verified | 64.5accuracy (%) | reproducedT1 | 2026-04-23Epoch AI |
| Chess Puzzles | 62.12accuracy (%) | reproducedT1 | 2026-04-23Epoch AI |
| CritPt | 30.57accuracy (%) | unverifiedT2 | 2026-04-23 |
Claims drawn from cited facts, not live model generation.
This model was developed by OpenAI and released in 2026.
2 cited facts
This model is scored on 10 tracked benchmarks and currently tops 2 of them. Because its results have been independently reproduced, the record can be treated as reliable.
3 cited facts
Under the disclosed harness, it trails the leader by 4.28 points on average, a moderate gap that places it in the competitive-but-not-front-of-pack band. It holds the top score on 2 benchmarks, which alongside the moderate average gap confirms competitive standing rather than category leadership.
3 cited facts
5 facts cross-checked across data sources: 3 corroborated, 1 single-source, 1 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: curated_capability_claims, epoch_benchmarks, litellm_prices, modelsdev_models +1 more
GPT-5.5 Pro is an AI model developed by OpenAI, released 2026-04-23. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-5.5 Pro has recorded scores on 10 benchmarks, each shown with its evidence status.
GPT-5.5 Pro has recorded scores on 10 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, Chess Puzzles, SimpleBench, SimpleQA Verified, CritPt, and 4 more. The full table above shows each score with its evidence status.
6 of 10 recorded scores (60%) are independently reproduced rather than self-reported by the lab.
GPT-5.5 Pro has 10 tracked claims: 6 independently reproduced, 4 unverified.
Listed API pricing: $30 per million input tokens, $180 per million output tokens. See the pricing block for the full breakdown.