OpenAI·released 2026-09-031 source
GPT-6 Astra benchmark scores: 16 benchmarks tracked, leading 11. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 11 of 16 benchmarks·best result 100 on OTIS Mock AIME 2024-2025 (unverified)·0 of 16 independently reproduced
Head-to-headGPT-6 Astra vs Gemini 3.1 ProLeads on: ARC-AGI, ARC-AGI-2, Chess Puzzles, DeepSWE, EBR-bench, FrontierMath-Tier-4-v2-Private, FrontierMath-Tiers-1-3-v2-Private, GPQA diamond, Mystery Game Puzzles, OTIS Mock AIME 2024-2025, SimpleQA Verified
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 100accuracy (%) | unverifiedT2 | 2026-09-03 |
| ProofBench | 99accuracy (%) | unverifiedT1 | 2026-09-03https://www.vals.ai/benchmarks/proof_bench |
| ARC-AGI | 98.5accuracy (%) | unverifiedT1 | 2026-09-03https://arcprize.org/leaderboard |
| FrontierMath-Tier-4-v2-Private | 97.6accuracy (%) | unverifiedT2 | 2026-09-03 |
| ARC-AGI-2 | 95accuracy (%) | unverifiedT2 | 2026-09-03 |
| GPQA diamond | 94.36accuracy (%) | unverifiedT2 | 2026-09-03 |
| FrontierMath-Tiers-1-3-v2-Private | 93.68accuracy (%) | unverifiedT2 | 2026-09-03 |
| WeirdML | 92.87accuracy (%) | unverifiedT1 | 2026-09-03https://htihle.github.io/weirdml.html |
| Mystery Game Puzzles | 82.37accuracy (%) | unverifiedT2 | 2026-09-03 |
| EBR-bench | 76.19accuracy (%) | unverifiedT2 | 2026-09-03 |
| SimpleQA Verified | 75.6accuracy (%) | unverifiedT2 | 2026-09-03 |
| DeepSWE | 74.12accuracy (%) | unverifiedT1 | 2026-09-03https://deepswe.datacurve.ai/ |
| Chess Puzzles | 70.54accuracy (%) | unverifiedT2 | 2026-09-03 |
| FrontierCode | 53.26accuracy (%) | unverifiedT1 | 2026-09-03https://cognition.com/frontiercode |
| MirrorCode | 46.7accuracy (%) | unverifiedT2 | 2026-09-03 |
| APEX-Agents | 46.7accuracy (%) | unverifiedT2 | 2026-09-03 |
Claims drawn from cited facts, not live model generation.
This model was developed by OpenAI. It dates to a 2026 vintage, making it a recent entry in its lineage, and no parameter count is offered to pin down its size.
2 cited facts
This model is scored on 16 tracked benchmarks and currently holds the top score on 11 of them. The record is unverified: outside evaluation has neither confirmed nor disputed these results, so they should be read as not yet independently validated.
3 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
GPT-6 Astra is an AI model developed by OpenAI, released 2026-09-03. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-6 Astra has recorded scores on 16 benchmarks, each shown with its evidence status.
GPT-6 Astra has recorded scores on 16 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, SimpleQA Verified, Chess Puzzles, ARC-AGI, and 10 more. The full table above shows each score with its evidence status.
0 of 16 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
GPT-6 Astra has 16 tracked claims: 16 unverified.