Mistral AI·released 2025-06-101 source
Magistral Small 1.0 benchmark scores: 4 benchmarks tracked. 50% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 4 benchmarks·best result 41.41 on GPQA diamond (reproduced)·2 of 4 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| GPQA diamond | 41.41accuracy (%) | reproduced· optimizedT1 | 2025-06-10Epoch AI |
| OTIS Mock AIME 2024-2025 | 29.93accuracy (%) | reproduced· optimizedT1 | 2025-06-10Epoch AI |
| ARC-AGI | 5accuracy (%) | unverified· optimizedT2 | 2025-06-10 |
| ARC-AGI-2 | 0accuracy (%) | unverifiedT2 | 2025-06-10 |
Claims drawn from cited facts, not live model generation.
This model was developed by Mistral AI and released in 2025, with no official parameter count disclosed for scale classification.
2 cited facts
This model is scored on four tracked benchmarks, but it holds no current top score in any of them. Because these results have been independently reproduced by outside evaluation, they can be treated as verified rather than vendor claims.
3 cited facts
With an average gap of 76.94 points, this model trails the leader by a wide margin on the disclosed harness, placing it well behind the front of the pack. It holds the top score on none of the benchmarks, so it has no category leadership yet.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Magistral Small 1.0 is an AI model developed by Mistral AI, released 2025-06-10. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Magistral Small 1.0 has recorded scores on 4 benchmarks, each shown with its evidence status.
Magistral Small 1.0 has recorded scores on 4 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, ARC-AGI, ARC-AGI-2. The full table above shows each score with its evidence status.
2 of 4 recorded scores (50%) are independently reproduced rather than self-reported by the lab.
Magistral Small 1.0 has 4 tracked claims: 2 independently reproduced, 2 unverified.