xAI·released 2025-02-191 source
Grok-3 mini benchmark scores: 11 benchmarks tracked. 45% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 11 benchmarks·best result 90.94 on MATH level 5 (reproduced)·5 of 11 independently reproduced·$0.3/$0.5 per M tokens
Consensus: LiteLLM · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MATH level 5 | 90.94accuracy (%) | reproduced· optimizedT1 | 2025-02-19Epoch AI |
| OTIS Mock AIME 2024-2025 | 77.76accuracy (%) | reproduced· optimizedT1 | 2025-02-19Epoch AI |
| Lech Mazur Writing | 73.5accuracy (%) | unverifiedT1 | 2025-02-19lechmazur/writing Github repository |
| GPQA diamond | 68.35accuracy (%) | reproduced· optimizedT1 | 2025-02-19Epoch AI |
| Fiction.LiveBench | 66.7accuracy (%) | unverifiedT1 | 2025-02-19Fiction.live leaderboard |
| Aider polyglot | 49.3accuracy (%) | unverified· optimizedT1 | 2025-02-19Aider LLM Leaderboards |
| WeirdML | 42.58accuracy (%) | unverifiedT1 | 2025-02-19WeirdML Leaderboard |
| SimpleQA Verified | 21.1accuracy (%) | reproducedT1 | 2025-02-19Epoch AI |
| ARC-AGI | 16.5accuracy (%) | unverified· optimizedT1 | 2025-02-19ARC Prize Leaderboard |
| FrontierMath-2025-02-28-Private | 10.28accuracy (%) | reproduced· optimizedT1 | 2025-02-19Epoch AI |
| ARC-AGI-2 | 0.42accuracy (%) | unverifiedT2 | 2025-02-19 |
Claims drawn from cited facts, not live model generation.
This model was developed by xAI and released in 2025.
2 cited facts
Enough cells here — 11 — that the shape of the record matters more than any single score. Absence of a lead is a statement about the top of each table only; it leaves the model's actual placement unstated. Because the scores have been independently reproduced, outside evaluation backs them up, so the record can be trusted.
3 cited facts
This model trails the state-of-the-art by an average of 42.05 points, a wide gap that places it well behind the leading models.
1 cited fact
6 facts cross-checked across data sources: 6 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices
Grok-3 mini is an AI model developed by xAI, released 2025-02-19. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Grok-3 mini has recorded scores on 11 benchmarks, each shown with its evidence status.
Grok-3 mini has recorded scores on 11 benchmarks — Lech Mazur Writing, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, FrontierMath-2025-02-28-Private, and 5 more. The full table above shows each score with its evidence status.
5 of 11 recorded scores (45%) are independently reproduced rather than self-reported by the lab.
Grok-3 mini has 11 tracked claims: 5 independently reproduced, 6 unverified.
Listed API pricing: $0.3 per million input tokens, $0.5 per million output tokens. See the pricing block for the full breakdown.