Mistral AI·released 2025-01-301 source
Mistral Small 3 benchmark scores: 4 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 4 benchmarks·best result 44.82 on MATH level 5 (unverified)·0 of 4 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MATH level 5 | 44.82accuracy (%) | unverifiedT2 | 2025-01-25 |
| GPQA diamond | 29.71accuracy (%) | unverifiedT2 | 2025-01-30 |
| OTIS Mock AIME 2024-2025 | 6.57accuracy (%) | unverifiedT2 | 2025-01-30 |
| Chess Puzzles | 0accuracy (%) | unverifiedT2 | 2025-01-30 |
Claims drawn from cited facts, not live model generation.
This model was developed by Mistral AI, a European AI lab founded to build efficient open-weight language models. It dates to 2025, placing it within the current generation of model releases.
2 cited facts
This model is scored on 4 tracked benchmarks and currently holds no top score among them. Its record remains unverified, meaning it is neither confirmed nor disputed by outside evaluation.
3 cited facts
Across its benchmarks the model trails the state-of-the-art leader by an average of 70.48 points, a wide gap that reads as well behind the front of the pack rather than merely a couple of points off the frontier. It also holds the top score on none of those benchmarks, meaning it shows no genuine category leadership on this evidence.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Mistral Small 3 is an AI model developed by Mistral AI, released 2025-01-30. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Mistral Small 3 has recorded scores on 4 benchmarks, each shown with its evidence status.
Mistral Small 3 has recorded scores on 4 benchmarks — GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, Chess Puzzles. The full table above shows each score with its evidence status.
0 of 4 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Mistral Small 3 has 4 tracked claims: 4 unverified.