OpenAI·released 2026-07-09✓ 2 sources
GPT-5.6 Sol benchmark scores: 20 benchmarks tracked, leading 5. 45% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 5 of 20 benchmarks·best result 100 on OTIS Mock AIME 2024-2025 (reproduced)·9 of 20 independently reproduced·$5/$30 per M tokens
Head-to-headGPT-5.6 Sol vs Claude Opus 5Consensus: LiteLLM · See every model’s pricing →
Leads on: ARC-AGI-2, Chess Puzzles, CritPt, FrontierMath-Tiers-1-3-v2-Private, OTIS Mock AIME 2024-2025
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 100accuracy (%) | reproduced· optimizedT1 | 2026-07-09Epoch AI |
| ARC-AGI | 97.5accuracy (%) | unverified· optimizedT1 | 2026-07-09https://arcprize.org/leaderboard |
| Surface Evolver Bench | 93.12accuracy (%) | unverifiedT1 | 2026-07-09https://yhenon.github.io/surface-evolver-llm-eval/ |
| ARC-AGI-2 | 92.5accuracy (%) | unverifiedT2 | 2026-07-09 |
| GPQA diamond | 91.33accuracy (%) | reproduced· optimizedT1 | 2026-07-09Epoch AI |
| WeirdML | 89.43accuracy (%) | unverifiedT1 | 2026-07-09https://htihle.github.io/weirdml.html |
| FrontierMath-Tiers-1-3-v2-Private | 89.12accuracy (%) | reproducedT1 | 2026-07-09Epoch AI |
| ProofBench | 83accuracy (%) | unverifiedT1 | 2026-07-09https://www.vals.ai/benchmarks/proof_bench |
| FrontierMath-Tier-4-v2-Private | 82.93accuracy (%) | reproducedT1 | 2026-07-09Epoch AI |
| DeepSWE | 72.67accuracy (%) | unverifiedT1 | 2026-07-09https://deepswe.datacurve.ai/ |
| SimpleQA Verified | 71.6accuracy (%) | reproducedT1 | 2026-07-09Epoch AI |
| SimpleBench | 66.04accuracy (%) | unverifiedT1 | 2026-07-09SimpleBench Leaderboard |
| Chess Puzzles | 62.12accuracy (%) | reproducedT1 | 2026-07-09Epoch AI |
| Mystery Game Puzzles | 53.73accuracy (%) | reproducedT1 | 2026-07-09Epoch AI |
| FrontierCode | 47.5accuracy (%) | unverifiedT1 | 2026-07-09https://cognition.com/frontiercode |
| APEX-Agents | 40accuracy (%) | unverifiedT2 | 2026-07-09 |
| EBR-bench | 39.05accuracy (%) | reproducedT1 | 2026-07-09Epoch AI |
| PostTrainBench | 36.23accuracy (%) | unverifiedT2 | 2026-07-09 |
| CritPt | 32.3accuracy (%) | unverifiedT2 | 2026-07-09 |
| MirrorCode | 20accuracy (%) | reproducedT1 | 2026-07-09Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by OpenAI and was released in 2026.
2 cited facts
This model is scored on 20 tracked benchmarks. It currently holds the top score on 5 of those benchmarks. Because the record is independently reproduced, outside evaluation backs up the scores and they can be treated as verified.
3 cited facts
Against disclosed harness scores, the model trails the state of the art by an average of 5.97 points, a moderate gap that reads as competitive but not front-of-pack. It nonetheless holds the top score on 5 benchmarks, indicating genuine category leadership on those tasks even though it is not the overall frontier leader.
2 cited facts
6 facts cross-checked across data sources: 1 corroborated, 5 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: curated_models, epoch_benchmarks, litellm_prices
GPT-5.6 Sol is an AI model developed by OpenAI, released 2026-07-09. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-5.6 Sol has recorded scores on 20 benchmarks, each shown with its evidence status.
GPT-5.6 Sol has recorded scores on 20 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleBench, SimpleQA Verified, and 14 more. The full table above shows each score with its evidence status.
9 of 20 recorded scores (45%) are independently reproduced rather than self-reported by the lab.
GPT-5.6 Sol has 20 tracked claims: 9 independently reproduced, 11 unverified.
Listed API pricing: $5 per million input tokens, $30 per million output tokens. See the pricing block for the full breakdown.