xAI·released 2026-02-171 source
Grok 4.20 benchmark scores: 13 benchmarks tracked. 46% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 13 benchmarks·best result 92.21 on OTIS Mock AIME 2024-2025 (reproduced)·6 of 13 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 92.21accuracy (%) | reproduced· optimizedT1 | 2026-02-17Epoch AI |
| ARC-AGI | 89.5accuracy (%) | unverified· optimizedT1 | 2026-02-17https://arcprize.org/leaderboard |
| GPQA diamond | 85.77accuracy (%) | reproduced· optimizedT1 | 2026-02-17Epoch AI |
| ARC-AGI-2 | 65.14accuracy (%) | unverifiedT2 | 2026-02-17 |
| Terminal Bench | 57.3accuracy (%) | unverified· optimizedT1 | 2026-02-17https://www.tbench.ai/leaderboard/terminal-bench/2.0 |
| WeirdML | 52.26accuracy (%) | unverifiedT1 | 2026-02-17https://htihle.github.io/weirdml.html |
| FrontierMath-Tiers-1-3-v2-Private | 44.91accuracy (%) | reproducedT1 | 2026-02-17Epoch AI |
| SimpleQA Verified | 35.1accuracy (%) | reproducedT1 | 2026-02-17Epoch AI |
| CL-bench | 22.2accuracy (%) | unverifiedT2 | 2026-02-17 |
| Chess Puzzles | 20.03accuracy (%) | reproducedT1 | 2026-02-17Epoch AI |
| FrontierMath-Tier-4-v2-Private | 17.07accuracy (%) | reproducedT1 | 2026-02-17Epoch AI |
| ProofBench | 14accuracy (%) | unverifiedT1 | 2026-02-17https://www.vals.ai/benchmarks/proof_bench |
| CL-bench Life | 11.9accuracy (%) | unverifiedT2 | 2026-02-17 |
Claims drawn from cited facts, not live model generation.
Developed by xAI, this model was released in 2026.
2 cited facts
This model is tracked across 13 benchmarks and currently holds no top score on any of them. Because the results are independently reproduced, readers can treat the record as verified rather than vendor-claimed.
3 cited facts
On average, this model trails the state-of-the-art leader by 32.21 points on its evaluated benchmarks (under the disclosed harness), a wide gap that places it well behind the front-of-pack. It holds the top score on none of the sota leader count benchmarks, so it has no genuine category leadership yet.
2 cited facts
5 facts cross-checked across data sources: 5 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, openrouter_models
Grok 4.20 is an AI model developed by xAI, released 2026-02-17. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Grok 4.20 has recorded scores on 13 benchmarks, each shown with its evidence status.
Grok 4.20 has recorded scores on 13 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleQA Verified, ARC-AGI, and 7 more. The full table above shows each score with its evidence status.
6 of 13 recorded scores (46%) are independently reproduced rather than self-reported by the lab.
Grok 4.20 has 13 tracked claims: 6 independently reproduced, 7 unverified.