Mistral AI·released 2025-09-181 source
Magistral Small 1.2 benchmark scores: 4 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 4 benchmarks·best result 35.52 on DTBench (unverified)·0 of 4 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| DTBench | 35.52accuracy (%) | unverifiedT1 | 2025-09-18https://conceptualreasoning.ai/dtbench |
| GPQA diamond | 30.13accuracy (%) | unverifiedT2 | 2025-09-18 |
| OTIS Mock AIME 2024-2025 | 27.98accuracy (%) | unverifiedT2 | 2025-09-18 |
| Chess Puzzles | 0accuracy (%) | unverifiedT2 | 2025-09-18 |
Claims drawn from cited facts, not live model generation.
This model was developed by Mistral AI, making it the work of an independent European lab rather than one of the large American or Chinese incumbents. It was released in 2025, which places it among the most recent generation of models rather than among older, more established architectures.
2 cited facts
This model is scored on four tracked benchmarks. It holds no current top score on any of them. Its record is unverified: outside evaluation has neither confirmed nor disputed the results, so they should be read as neither validated nor contested.
3 cited facts
The model trails the state of the art by 67.15 points on average, a wide gap under the disclosed harness that reads as well behind the leaders rather than at the frontier or competitive. It holds the top score on 0 benchmarks, meaning it currently shows no category leadership.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Magistral Small 1.2 is an AI model developed by Mistral AI, released 2025-09-18. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Magistral Small 1.2 has recorded scores on 4 benchmarks, each shown with its evidence status.
Magistral Small 1.2 has recorded scores on 4 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, DTBench, Chess Puzzles. The full table above shows each score with its evidence status.
0 of 4 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Magistral Small 1.2 has 4 tracked claims: 4 unverified.