xAI·released 2025-07-091 source
Grok 4 benchmark scores: 21 benchmarks tracked. 29% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 21 benchmarks·best result 94.4 on Fiction.LiveBench (unverified)·6 of 21 independently reproduced·$3/$15 per M tokens
Consensus: LiteLLM · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| Fiction.LiveBench | 94.4accuracy (%) | unverifiedT1 | 2025-07-09Fiction.live leaderboard |
| OTIS Mock AIME 2024-2025 | 83.98accuracy (%) | reproduced· optimizedT1 | 2025-07-09Epoch AI |
| GPQA diamond | 82.67accuracy (%) | reproduced· optimizedT1 | 2025-07-09Epoch AI |
| Aider polyglot | 79.6accuracy (%) | unverified· optimizedT1 | 2025-07-09Aider LLM Leaderboards |
| Lech Mazur Writing | 76.9accuracy (%) | unverifiedT1 | 2025-07-09lechmazur/writing Github repository |
| ARC-AGI | 66.67accuracy (%) | unverified· optimizedT2 | 2025-07-09 |
| METR Time Horizons | 66.58accuracy (%) | unverifiedT1 | 2025-07-09METR - Measuring AI Ability to Complete Long Tasks |
| SimpleBench | 52.6accuracy (%) | unverifiedT1 | 2025-07-09SimpleBench Leaderboard |
| SimpleQA Verified | 47.9accuracy (%) | reproducedT1 | 2025-07-09Epoch AI |
| DeepResearch Bench | 47.9accuracy (%) | unverifiedT1 | 2025-07-09DeepResearchBench Leaderboard |
| WeirdML | 45.73accuracy (%) | unverifiedT1 | 2025-07-09WeirdML Leaderboard |
| GeoBench | 45accuracy (%) | unverifiedT1 | 2025-07-09GeoBench leaderboard |
| Balrog | 43.6accuracy (%) | unverifiedT1 | 2025-07-09Balrog Leaderboard |
| Cybench | 43accuracy (%) | self-reported· optimizedT1 | 2025-07-09Grok 4 Model Card |
| FrontierMath-2025-02-28-Private | 34.48accuracy (%) | reproduced· optimizedT1 | 2025-07-09Epoch AI |
| Terminal Bench | 27.2accuracy (%) | unverified· optimizedT1 | 2025-07-09Terminal-Bench v2 Leaderboard |
| Chess Puzzles | 24.24accuracy (%) | reproducedT1 | 2025-07-09Epoch AI |
| GDPval | 21.1accuracy (%) | unverifiedT2 | 2025-07-09 |
| ARC-AGI-2 | 15.97accuracy (%) | unverifiedT2 | 2025-07-09 |
| APEX-Agents | 15.2accuracy (%) | unverifiedT2 | 2025-07-09 |
| FrontierMath-Tier-4-2025-07-01-Private | 3.47accuracy (%) | reproducedT1 | 2025-07-09Epoch AI |
Claims drawn from cited facts, not live model generation.
This model is scored on 21 tracked benchmarks and, among those, holds no current top score. Because the record is independently reproduced, outside evaluation backs the scores, so readers can treat them as verified results rather than vendor claims.
3 cited facts
With an average gap of 28.53 points to the state of the art, this model sits well behind the leaders. It holds the top score on none of the tracked benchmarks, meaning no category leadership yet.
2 cited facts
6 facts cross-checked across data sources: 6 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices
Grok 4 is an AI model developed by xAI, released 2025-07-09. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Grok 4 has recorded scores on 21 benchmarks, each shown with its evidence status.
Grok 4 has recorded scores on 21 benchmarks — Lech Mazur Writing, GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, Cybench, and 15 more. The full table above shows each score with its evidence status.
6 of 21 recorded scores (29%) are independently reproduced rather than self-reported by the lab.
Grok 4 has 21 tracked claims: 6 independently reproduced, 1 self-reported, 14 unverified.
Listed API pricing: $3 per million input tokens, $15 per million output tokens. See the pricing block for the full breakdown.