OpenAI·released 2026-05-051 source
GPT-5.5 Instant benchmark scores: 7 benchmarks tracked. 86% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 7 benchmarks·best result 76.68 on GPQA diamond (reproduced)·6 of 7 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| GPQA diamond | 76.68accuracy (%) | reproduced· optimizedT1 | 2026-05-05Epoch AI |
| OTIS Mock AIME 2024-2025 | 68.02accuracy (%) | reproduced· optimizedT1 | 2026-05-05Epoch AI |
| SimpleQA Verified | 38accuracy (%) | reproducedT1 | 2026-05-05Epoch AI |
| FrontierMath-Tiers-1-3-v2-Private | 26.32accuracy (%) | reproducedT1 | 2026-05-05Epoch AI |
| Chess Puzzles | 7.41accuracy (%) | reproducedT1 | 2026-05-05Epoch AI |
| FrontierMath-Tier-4-v2-Private | 2.44accuracy (%) | reproducedT1 | 2026-05-05Epoch AI |
| CritPt | 0accuracy (%) | unverifiedT2 | 2026-05-05 |
Claims drawn from cited facts, not live model generation.
Developed by OpenAI and released in 2026, this model has no disclosed parameter count, leaving its scale uncharacterized.
2 cited facts
This model is scored on seven tracked benchmarks and currently holds no top score in any of them. Because the underlying scores have been independently reproduced by outside evaluation, the record can be treated as verified rather than as vendor-claimed.
3 cited facts
With an average SOTA gap of 46.12 points, this model sits well behind the leaders on its benchmarks. It holds no top scores on the tracked leaderboards, so there is no current category leadership to report.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
GPT-5.5 Instant is an AI model developed by OpenAI, released 2026-05-05. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-5.5 Instant has recorded scores on 7 benchmarks, each shown with its evidence status.
GPT-5.5 Instant has recorded scores on 7 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, Chess Puzzles, SimpleQA Verified, CritPt, FrontierMath-Tiers-1-3-v2-Private, and 1 more. The full table above shows each score with its evidence status.
6 of 7 recorded scores (86%) are independently reproduced rather than self-reported by the lab.
GPT-5.5 Instant has 7 tracked claims: 6 independently reproduced, 1 unverified.