Mistral AI·released 2025-06-201 source
Mistral Small 3.2 benchmark scores: 4 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 4 benchmarks·best result 33.12 on DTBench (unverified)·0 of 4 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| DTBench | 33.12accuracy (%) | unverifiedT1 | 2025-06-20https://conceptualreasoning.ai/dtbench |
| GPQA diamond | 32.07accuracy (%) | unverifiedT2 | 2025-06-20 |
| OTIS Mock AIME 2024-2025 | 30.21accuracy (%) | unverifiedT2 | 2025-06-20 |
| Chess Puzzles | 0accuracy (%) | unverifiedT2 | 2025-06-20 |
Claims drawn from cited facts, not live model generation.
This model is developed by Mistral AI and was released in 2025, placing it among the more recent additions to the field rather than an older, established lineage. No parameter count is available for this model, so its scale cannot be characterized here.
2 cited facts
This model is scored on 4 tracked benchmarks, and it currently holds no top score among them. Its record is unverified, meaning neither confirmed nor disputed by outside evaluation.
3 cited facts
Across its benchmarks, this model trails the leading score by 66.71 points on average — a wide margin that reads as well behind the leaders rather than merely competitive, and it should be understood as a score gap under a disclosed harness, not a claim about capability. It holds the top score on none of its benchmarks (leader count of 0), so there is no category leadership here: the wide average gap and the zero leader count together place it clearly off the frontier.
3 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Mistral Small 3.2 is an AI model developed by Mistral AI, released 2025-06-20. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Mistral Small 3.2 has recorded scores on 4 benchmarks, each shown with its evidence status.
Mistral Small 3.2 has recorded scores on 4 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, DTBench, Chess Puzzles. The full table above shows each score with its evidence status.
0 of 4 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Mistral Small 3.2 has 4 tracked claims: 4 unverified.