xAI·released 2026-07-081 source
Grok 4.5 benchmark scores: 17 benchmarks tracked. 35% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 17 benchmarks·best result 97.78 on OTIS Mock AIME 2024-2025 (reproduced)·6 of 17 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 97.78accuracy (%) | reproduced· optimizedT1 | 2026-07-08Epoch AI |
| GPQA diamond | 91.25accuracy (%) | reproduced· optimizedT1 | 2026-07-08Epoch AI |
| ARC-AGI | 87.17accuracy (%) | unverified· optimizedT1 | 2026-07-08https://arcprize.org/leaderboard |
| Surface Evolver Bench | 74.38accuracy (%) | unverifiedT1 | 2026-07-08https://yhenon.github.io/surface-evolver-llm-eval/ |
| SimpleBench | 64accuracy (%) | unverifiedT1 | 2026-07-08SimpleBench Leaderboard |
| FrontierMath-Tiers-1-3-v2-Private | 57.19accuracy (%) | reproducedT1 | 2026-07-08Epoch AI |
| DeepSWE | 53.76accuracy (%) | unverifiedT1 | 2026-07-08https://deepswe.datacurve.ai/ |
| SimpleQA Verified | 53.5accuracy (%) | reproducedT1 | 2026-07-08Epoch AI |
| ARC-AGI-2 | 52.64accuracy (%) | unverifiedT2 | 2026-07-08 |
| WeirdML | 46.43accuracy (%) | unverifiedT1 | 2026-07-08https://htihle.github.io/weirdml.html |
| FrontierCode | 42.4accuracy (%) | unverifiedT1 | 2026-07-08https://cognition.com/frontiercode |
| APEX-Agents | 34.2accuracy (%) | unverifiedT2 | 2026-07-08 |
| Chess Puzzles | 32.66accuracy (%) | reproducedT1 | 2026-07-08Epoch AI |
| ProofBench | 31accuracy (%) | unverifiedT1 | 2026-07-08https://www.vals.ai/benchmarks/proof_bench |
| FrontierMath-Tier-4-v2-Private | 24.39accuracy (%) | reproducedT1 | 2026-07-08Epoch AI |
| PostTrainBench | 23.45accuracy (%) | unverifiedT2 | 2026-07-08 |
| CritPt | 15.43accuracy (%) | unverifiedT2 | 2026-07-08 |
Claims drawn from cited facts, not live model generation.
This model was developed by xAI and released in 2026.
2 cited facts
This model is scored on 17 tracked benchmarks and currently holds no top score in any of them. Because the results have been independently reproduced, the record can be trusted as verified rather than as vendor claims.
3 cited facts
On average this model trails the SOTA leader by 25.25 points, a wide gap that places it well behind the front of the pack. It currently holds the top score on none of the considered benchmarks.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Grok 4.5 is an AI model developed by xAI, released 2026-07-08. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Grok 4.5 has recorded scores on 17 benchmarks, each shown with its evidence status.
Grok 4.5 has recorded scores on 17 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleBench, SimpleQA Verified, and 11 more. The full table above shows each score with its evidence status.
6 of 17 recorded scores (35%) are independently reproduced rather than self-reported by the lab.
Grok 4.5 has 17 tracked claims: 6 independently reproduced, 11 unverified.