Mistral AI·released 2025-05-071 source
Mistral Medium 3 benchmark scores: 6 benchmarks tracked. 67% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 6 benchmarks·best result 81.63 on MATH level 5 (reproduced)·4 of 6 independently reproduced·$0.4/$2 per M tokens
Consensus: LiteLLM · Cross-check: OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MATH level 5 | 81.63accuracy (%) | reproduced· optimizedT1 | 2025-05-07Epoch AI |
| Lech Mazur Writing | 77.3accuracy (%) | unverifiedT1 | 2025-05-07lechmazur/writing Github repository |
| GPQA diamond | 46.04accuracy (%) | reproduced· optimizedT1 | 2025-05-07Epoch AI |
| OTIS Mock AIME 2024-2025 | 32.15accuracy (%) | reproduced· optimizedT1 | 2025-05-07Epoch AI |
| FrontierMath-2025-02-28-Private | 0.61accuracy (%) | reproduced· optimizedT1 | 2025-05-07Epoch AI |
| HLE | 0accuracy (%) | unverifiedT2 | 2025-05-07 |
Claims drawn from cited facts, not live model generation.
This model was developed by Mistral AI. It was released in 2025.
2 cited facts
Modest evidence: 6 tracked benchmarks without a lead orients the reader and settles nothing. The record is independently reproduced, meaning outside evaluation backs the scores up.
3 cited facts
On average, this model trails the state-of-the-art leader by 41.89 points, a wide gap that indicates it is well behind the leaders. It holds the top score on zero benchmarks, confirming it is not at the frontier.
2 cited facts
5 facts cross-checked across data sources: 2 corroborated, 1 single-source, 2 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, openrouter_models
Mistral Medium 3 is an AI model developed by Mistral AI, released 2025-05-07. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Mistral Medium 3 has recorded scores on 6 benchmarks, each shown with its evidence status.
Mistral Medium 3 has recorded scores on 6 benchmarks — Lech Mazur Writing, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, FrontierMath-2025-02-28-Private, HLE. The full table above shows each score with its evidence status.
4 of 6 recorded scores (67%) are independently reproduced rather than self-reported by the lab.
Mistral Medium 3 has 6 tracked claims: 4 independently reproduced, 2 unverified.
Listed API pricing: $0.4 per million input tokens, $2 per million output tokens (prices disputed across sources). See the pricing block for the full breakdown.