Mistral AI·released 2025-03-171 source
Mistral Small 3.1 benchmark scores: 4 benchmarks tracked. 75% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 4 benchmarks·best result 46.77 on MATH level 5 (reproduced)·3 of 4 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MATH level 5 | 46.77accuracy (%) | reproduced· optimizedT1 | 2025-03-17Epoch AI |
| GPQA diamond | 29.97accuracy (%) | reproduced· optimizedT1 | 2025-03-17Epoch AI |
| OTIS Mock AIME 2024-2025 | 5.74accuracy (%) | reproduced· optimizedT1 | 2025-03-17Epoch AI |
| CritPt | 0accuracy (%) | unverifiedT2 | 2025-03-17 |
Claims drawn from cited facts, not live model generation.
This model was developed by Mistral AI and released in 2025.
2 cited facts
This model is scored on 4 tracked benchmarks, currently does not top any of them, and, as the record is independently reproduced, the scores are independently verified and can be trusted.
3 cited facts
This model trails the leader by an average of 59.76 points, a wide gap that places it well behind the top-performing models. It does not hold the top score on any benchmark, confirming its position far from the frontier.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Mistral Small 3.1 is an AI model developed by Mistral AI, released 2025-03-17. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Mistral Small 3.1 has recorded scores on 4 benchmarks, each shown with its evidence status.
Mistral Small 3.1 has recorded scores on 4 benchmarks — GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, CritPt. The full table above shows each score with its evidence status.
3 of 4 recorded scores (75%) are independently reproduced rather than self-reported by the lab.
Mistral Small 3.1 has 4 tracked claims: 3 independently reproduced, 1 unverified.