xAI·released 2026-04-171 source
Grok 4.3 Beta benchmark scores: 9 benchmarks tracked. 67% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 9 benchmarks·best result 93.33 on OTIS Mock AIME 2024-2025 (reproduced)·6 of 9 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 93.33accuracy (%) | reproduced· optimizedT1 | 2026-04-17Epoch AI |
| GPQA diamond | 85.1accuracy (%) | reproduced· optimizedT1 | 2026-04-17Epoch AI |
| WeirdML | 49.89accuracy (%) | unverifiedT1 | 2026-04-17https://htihle.github.io/weirdml.html |
| FrontierMath-Tiers-1-3-v2-Private | 42.81accuracy (%) | reproducedT1 | 2026-04-17Epoch AI |
| SimpleQA Verified | 38accuracy (%) | reproducedT1 | 2026-04-17Epoch AI |
| Chess Puzzles | 21.09accuracy (%) | reproducedT1 | 2026-04-17Epoch AI |
| FrontierMath-Tier-4-v2-Private | 14.63accuracy (%) | reproducedT1 | 2026-04-17Epoch AI |
| ProofBench | 11accuracy (%) | unverifiedT1 | 2026-04-17https://www.vals.ai/benchmarks/proof_bench |
| CritPt | 8accuracy (%) | unverifiedT2 | 2026-04-17 |
Claims drawn from cited facts, not live model generation.
This model, developed by xAI, was released in 2026.
2 cited facts
This model is scored on 9 tracked benchmarks and holds no current top score on any of them. Because outside evaluation independently reproduces the reported results, this benchmark record can be treated as verified rather than merely vendor-claimed.
3 cited facts
With an average gap of 40.98 points behind the leader, this model sits well behind the frontier on its benchmarks. It holds the top score on none of them, meaning it has no category leadership yet.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Grok 4.3 Beta is an AI model developed by xAI, released 2026-04-17. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Grok 4.3 Beta has recorded scores on 9 benchmarks, each shown with its evidence status.
Grok 4.3 Beta has recorded scores on 9 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleQA Verified, CritPt, and 3 more. The full table above shows each score with its evidence status.
6 of 9 recorded scores (67%) are independently reproduced rather than self-reported by the lab.
Grok 4.3 Beta has 9 tracked claims: 6 independently reproduced, 3 unverified.