Mistral AI·released 2024-07-241 source
Mistral Large 2 (Jul 2024) benchmark scores: 6 benchmarks tracked. 50% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 6 benchmarks·best result 73.33 on MMLU (unverified)·3 of 6 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MMLU | 73.33accuracy (%) | unverified· optimizedT1 | 2024-07-24Stanford CRFM Leaderboard |
| Lech Mazur Writing | 69accuracy (%) | unverifiedT1 | 2024-07-24lechmazur/writing Github repository |
| MATH level 5 | 44.82accuracy (%) | reproduced· optimizedT1 | 2024-07-24Epoch AI |
| GPQA diamond | 32.03accuracy (%) | reproduced· optimizedT1 | 2024-07-24Epoch AI |
| OTIS Mock AIME 2024-2025 | 8.38accuracy (%) | reproduced· optimizedT1 | 2024-07-24Epoch AI |
| SimpleBench | 7accuracy (%) | unverifiedT1 | 2024-07-24SimpleBench Leaderboard |
Claims drawn from cited facts, not live model generation.
This model was developed by Mistral AI and released in 2024.
2 cited facts
Coverage stands at 6 tracked benchmarks — thin enough that each new score will still move the overall picture. Leading none of them at present bounds the record's ceiling, not the model — and the reading is dated the moment new scores land. These scores come from independent reproduction, so the record is trustworthy.
3 cited facts
The model trails the state-of-the-art by an average of 51.01 points across benchmarks, placing it well behind the leaders.
1 cited fact
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Mistral Large 2 (Jul 2024) is an AI model developed by Mistral AI, released 2024-07-24. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Mistral Large 2 (Jul 2024) has recorded scores on 6 benchmarks, each shown with its evidence status.
Mistral Large 2 (Jul 2024) has recorded scores on 6 benchmarks — Lech Mazur Writing, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, MMLU, SimpleBench. The full table above shows each score with its evidence status.
3 of 6 recorded scores (50%) are independently reproduced rather than self-reported by the lab.
Mistral Large 2 (Jul 2024) has 6 tracked claims: 3 independently reproduced, 3 unverified.