OpenAI·released 2026-07-09✓ 2 sources
GPT-5.6 Terra benchmark scores: 16 benchmarks tracked. 44% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 16 benchmarks·best result 99.72 on OTIS Mock AIME 2024-2025 (reproduced)·7 of 16 independently reproduced·$2/$12 per M tokens
Consensus: LiteLLM · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 99.72accuracy (%) | reproduced· optimizedT1 | 2026-07-09Epoch AI |
| ARC-AGI | 96.5accuracy (%) | unverified· optimizedT1 | 2026-07-09https://arcprize.org/leaderboard |
| GPQA diamond | 91.08accuracy (%) | reproduced· optimizedT1 | 2026-07-09Epoch AI |
| FrontierMath-Tiers-1-3-v2-Private | 85.96accuracy (%) | reproducedT1 | 2026-07-09Epoch AI |
| ARC-AGI-2 | 83.9accuracy (%) | unverifiedT2 | 2026-07-09 |
| Surface Evolver Bench | 83.75accuracy (%) | unverifiedT1 | 2026-07-09https://yhenon.github.io/surface-evolver-llm-eval/ |
| WeirdML | 78.27accuracy (%) | unverifiedT1 | 2026-07-09https://htihle.github.io/weirdml.html |
| ProofBench | 71accuracy (%) | unverifiedT1 | 2026-07-09https://www.vals.ai/benchmarks/proof_bench |
| FrontierMath-Tier-4-v2-Private | 70.73accuracy (%) | reproducedT1 | 2026-07-09Epoch AI |
| DeepSWE | 69.62accuracy (%) | unverifiedT1 | 2026-07-09https://deepswe.datacurve.ai/ |
| Chess Puzzles | 51.6accuracy (%) | reproducedT1 | 2026-07-09Epoch AI |
| SimpleQA Verified | 43.1accuracy (%) | reproducedT1 | 2026-07-09Epoch AI |
| FrontierCode | 41.3accuracy (%) | unverifiedT1 | 2026-07-09https://cognition.com/frontiercode |
| SimpleBench | 38.68accuracy (%) | unverifiedT1 | 2026-07-09SimpleBench Leaderboard |
| CritPt | 30accuracy (%) | unverifiedT2 | 2026-07-09 |
| Mystery Game Puzzles | 28.4accuracy (%) | reproducedT1 | 2026-07-09Epoch AI |
Claims drawn from cited facts, not live model generation.
This model, developed by OpenAI, was released in 2026; no parameter count is offered, so its scale remains unspecified.
2 cited facts
This model is tracked across 16 benchmarks, but it holds no current top score on any of them. Because its record is independently reproduced, these results are backed by outside evaluation and can be treated as verified.
3 cited facts
Relative to the disclosed harness, it trails the best score by an average of 13.46 points, a wide gap that places it well behind the front of the pack. It holds the top score on 0 benchmarks, so there is no category leadership yet.
2 cited facts
6 facts cross-checked across data sources: 1 corroborated, 5 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: curated_models, epoch_benchmarks, litellm_prices
GPT-5.6 Terra is an AI model developed by OpenAI, released 2026-07-09. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-5.6 Terra has recorded scores on 16 benchmarks, each shown with its evidence status.
GPT-5.6 Terra has recorded scores on 16 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleBench, SimpleQA Verified, and 10 more. The full table above shows each score with its evidence status.
7 of 16 recorded scores (44%) are independently reproduced rather than self-reported by the lab.
GPT-5.6 Terra has 16 tracked claims: 7 independently reproduced, 9 unverified.
Listed API pricing: $2 per million input tokens, $12 per million output tokens. See the pricing block for the full breakdown.