xAI·released 2024-08-131 source
Grok-2 (Dec 2024) benchmark scores: 7 benchmarks tracked. 57% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 7 benchmarks·best result 63.6 on Lech Mazur Writing (unverified)·4 of 7 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| Lech Mazur Writing | 63.6accuracy (%) | unverifiedT1 | 2024-08-13lechmazur/writing Github repository |
| MATH level 5 | 63.52accuracy (%) | reproduced· optimizedT1 | 2024-08-13Epoch AI |
| GPQA diamond | 38.38accuracy (%) | reproduced· optimizedT1 | 2024-08-13Epoch AI |
| WeirdML | 22.24accuracy (%) | unverifiedT1 | 2024-08-13WeirdML Leaderboard |
| OTIS Mock AIME 2024-2025 | 11.44accuracy (%) | reproduced· optimizedT1 | 2024-08-13Epoch AI |
| SimpleBench | 7.24accuracy (%) | unverifiedT1 | 2024-08-13SimpleBench Leaderboard |
| FrontierMath-2025-02-28-Private | 1.21accuracy (%) | reproduced· optimizedT1 | 2024-08-13Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by xAI. It was released in 2024.
2 cited facts
Modest coverage here: 7 tracked scores and none of them a lead — too little to rank the model and no reason to mark it down. Because the scores have been independently reproduced, the record can be trusted as verified.
3 cited facts
This model trails the leader by an average of 57.88 points, a wide gap that places it well behind the front of the pack. It holds the top score on none of the benchmarks, indicating no category leadership.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Grok-2 (Dec 2024) is an AI model developed by xAI, released 2024-08-13. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Grok-2 (Dec 2024) has recorded scores on 7 benchmarks, each shown with its evidence status.
Grok-2 (Dec 2024) has recorded scores on 7 benchmarks — Lech Mazur Writing, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, SimpleBench, and 1 more. The full table above shows each score with its evidence status.
4 of 7 recorded scores (57%) are independently reproduced rather than self-reported by the lab.
Grok-2 (Dec 2024) has 7 tracked claims: 4 independently reproduced, 3 unverified.