xAI·released 2026-08-121 source
Grok 4.6 benchmark scores: 14 benchmarks tracked. 50% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 14 benchmarks·best result 99.17 on OTIS Mock AIME 2024-2025 (reproduced)·7 of 14 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 99.17accuracy (%) | reproduced· optimizedT1 | 2026-08-12Epoch AI |
| GPQA diamond | 92accuracy (%) | reproduced· optimizedT1 | 2026-08-12Epoch AI |
| ARC-AGI | 87.5accuracy (%) | unverified· optimizedT1 | 2026-08-12https://arcprize.org/leaderboard |
| DeepSWE | 67.48accuracy (%) | unverifiedT1 | 2026-08-12https://deepswe.datacurve.ai/ |
| WeirdML | 67.29accuracy (%) | unverifiedT1 | 2026-08-12https://htihle.github.io/weirdml.html |
| ARC-AGI-2 | 67.08accuracy (%) | unverifiedT2 | 2026-08-12 |
| FrontierMath-Tiers-1-3-v2-Private | 65.96accuracy (%) | reproducedT1 | 2026-08-12Epoch AI |
| SimpleQA Verified | 54.4accuracy (%) | reproducedT1 | 2026-08-12Epoch AI |
| ProofBench | 51accuracy (%) | unverifiedT1 | 2026-08-12https://www.vals.ai/benchmarks/proof_bench |
| FrontierCode | 48.01accuracy (%) | unverifiedT1 | 2026-08-12https://cognition.com/frontiercode |
| APEX-Agents | 41.2accuracy (%) | unverifiedT2 | 2026-08-12 |
| Chess Puzzles | 36.87accuracy (%) | reproducedT1 | 2026-08-12Epoch AI |
| FrontierMath-Tier-4-v2-Private | 31.71accuracy (%) | reproducedT1 | 2026-08-12Epoch AI |
| Mystery Game Puzzles | 27.3accuracy (%) | reproducedT1 | 2026-08-12Epoch AI |
Claims drawn from cited facts, not live model generation.
Developed by xAI, this model was released in 2026.
2 cited facts
This model is scored on 14 tracked benchmarks and currently holds no top score in any of them. Its results are independently reproduced, so outside evaluation backs the scores up and they can be treated as verified.
3 cited facts
With an average gap of 20.1 points behind the top score, the model sits well behind the leaders rather than at the frontier. It holds the top score on none of the tracked benchmarks, so it shows no category leadership yet.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Grok 4.6 is an AI model developed by xAI, released 2026-08-12. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Grok 4.6 has recorded scores on 14 benchmarks, each shown with its evidence status.
Grok 4.6 has recorded scores on 14 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleQA Verified, ARC-AGI, and 8 more. The full table above shows each score with its evidence status.
7 of 14 recorded scores (50%) are independently reproduced rather than self-reported by the lab.
Grok 4.6 has 14 tracked claims: 7 independently reproduced, 7 unverified.