OpenAI·released 2026-07-09✓ 2 sources
GPT-5.6 Luna benchmark scores: 16 benchmarks tracked. 44% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 16 benchmarks·best result 98.33 on OTIS Mock AIME 2024-2025 (reproduced)·7 of 16 independently reproduced·$0.2/$1.2 per M tokens
Consensus: LiteLLM · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 98.33accuracy (%) | reproduced· optimizedT1 | 2026-07-09Epoch AI |
| GPQA diamond | 88.8accuracy (%) | reproduced· optimizedT1 | 2026-07-09Epoch AI |
| ARC-AGI | 88accuracy (%) | unverified· optimizedT1 | 2026-07-09https://arcprize.org/leaderboard |
| FrontierMath-Tiers-1-3-v2-Private | 82.11accuracy (%) | reproducedT1 | 2026-07-09Epoch AI |
| DeepSWE | 67.19accuracy (%) | unverifiedT1 | 2026-07-09https://deepswe.datacurve.ai/ |
| FrontierMath-Tier-4-v2-Private | 60.98accuracy (%) | reproducedT1 | 2026-07-09Epoch AI |
| WeirdML | 60.86accuracy (%) | unverifiedT1 | 2026-07-09https://htihle.github.io/weirdml.html |
| Surface Evolver Bench | 60accuracy (%) | unverifiedT1 | 2026-07-09https://yhenon.github.io/surface-evolver-llm-eval/ |
| ProofBench | 60accuracy (%) | unverifiedT1 | 2026-07-09https://www.vals.ai/benchmarks/proof_bench |
| ARC-AGI-2 | 59.54accuracy (%) | unverifiedT2 | 2026-07-09 |
| SimpleQA Verified | 41.7accuracy (%) | reproducedT1 | 2026-07-09Epoch AI |
| FrontierCode | 39.8accuracy (%) | unverifiedT1 | 2026-07-09https://cognition.com/frontiercode |
| Chess Puzzles | 36.87accuracy (%) | reproducedT1 | 2026-07-09Epoch AI |
| SimpleBench | 36.16accuracy (%) | unverifiedT1 | 2026-07-09SimpleBench Leaderboard |
| CritPt | 20.6accuracy (%) | unverifiedT2 | 2026-07-09 |
| Mystery Game Puzzles | 12.98accuracy (%) | reproducedT1 | 2026-07-09Epoch AI |
Claims drawn from cited facts, not live model generation.
Developed by OpenAI and released in 2026, this model has no specified parameter count, so its scale-related trade-offs cannot be stated.
2 cited facts
This model is evaluated on 16 tracked benchmarks, yet it tops none of them. Because the scores have been independently reproduced, this record can be trusted as verified rather than treated as mere vendor claims.
3 cited facts
With an average SOTA gap of 22.81 points under the disclosed harness, this model trails the leader by a wide margin — well behind the front of the pack. It holds the top score on none of the tracked benchmarks, so there is no category leadership yet.
2 cited facts
6 facts cross-checked across data sources: 1 corroborated, 5 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: curated_models, epoch_benchmarks, litellm_prices
GPT-5.6 Luna is an AI model developed by OpenAI, released 2026-07-09. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-5.6 Luna has recorded scores on 16 benchmarks, each shown with its evidence status.
GPT-5.6 Luna has recorded scores on 16 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleBench, SimpleQA Verified, and 10 more. The full table above shows each score with its evidence status.
7 of 16 recorded scores (44%) are independently reproduced rather than self-reported by the lab.
GPT-5.6 Luna has 16 tracked claims: 7 independently reproduced, 9 unverified.
Listed API pricing: $0.2 per million input tokens, $1.2 per million output tokens. See the pricing block for the full breakdown.