xAI·released 2025-02-171 source
Grok 3 benchmark scores: 14 benchmarks tracked. 36% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 14 benchmarks·best result 88.75 on MATH level 5 (reproduced)·5 of 14 independently reproduced·$3/$15 per M tokens
Consensus: LiteLLM · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MATH level 5 | 88.75accuracy (%) | reproduced· optimizedT1 | 2025-02-17Epoch AI |
| Lech Mazur Writing | 76.4accuracy (%) | unverifiedT1 | 2025-02-17lechmazur/writing Github repository |
| GPQA diamond | 67.68accuracy (%) | reproduced· optimizedT1 | 2025-02-17Epoch AI |
| Fiction.LiveBench | 58.3accuracy (%) | unverifiedT1 | 2025-02-17Fiction.live leaderboard |
| OTIS Mock AIME 2024-2025 | 55.51accuracy (%) | reproduced· optimizedT1 | 2025-02-17Epoch AI |
| Aider polyglot | 53.3accuracy (%) | unverified· optimizedT1 | 2025-02-17Aider LLM Leaderboards |
| WeirdML | 37.24accuracy (%) | unverifiedT1 | 2025-02-17WeirdML Leaderboard |
| Balrog | 29.5accuracy (%) | unverifiedT1 | 2025-02-17Balrog Leaderboard |
| SimpleBench | 23.32accuracy (%) | unverifiedT1 | 2025-02-17SimpleBench Leaderboard |
| FrontierMath-2025-02-28-Private | 6.65accuracy (%) | reproduced· optimizedT1 | 2025-02-17Epoch AI |
| ARC-AGI | 5.5accuracy (%) | unverified· optimizedT1 | 2025-02-17ARC Prize Leaderboard |
| APEX-Agents | 2.1accuracy (%) | unverifiedT2 | 2025-02-17 |
| FrontierMath-Tier-4-2025-07-01-Private | 0accuracy (%) | reproducedT1 | 2025-02-17Epoch AI |
| ARC-AGI-2 | 0accuracy (%) | unverifiedT2 | 2025-02-17 |
Claims drawn from cited facts, not live model generation.
This model was developed by xAI and released in 2025.
2 cited facts
Breadth without a front-running result — 14 scored benchmarks and no leads — supports pattern-reading across tables while claiming nothing about overall standing. The scores have been independently reproduced by external evaluation, so the reader can trust the record.
3 cited facts
On average, the model trails the state-of-the-art leader by 43.98 points, a wide gap that places it well behind the frontier. It holds the top score on none of the benchmarks, indicating no category leadership.
2 cited facts
6 facts cross-checked across data sources: 6 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices
Grok 3 is an AI model developed by xAI, released 2025-02-17. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Grok 3 has recorded scores on 14 benchmarks, each shown with its evidence status.
Grok 3 has recorded scores on 14 benchmarks — Lech Mazur Writing, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, SimpleBench, and 8 more. The full table above shows each score with its evidence status.
5 of 14 recorded scores (36%) are independently reproduced rather than self-reported by the lab.
Grok 3 has 14 tracked claims: 5 independently reproduced, 9 unverified.
Listed API pricing: $3 per million input tokens, $15 per million output tokens. See the pricing block for the full breakdown.