Mistral AI·released 2024-02-261 source
Mistral Large benchmark scores: 4 benchmarks tracked. 75% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 4 benchmarks·best result 58.4 on MMLU (unverified)·3 of 4 independently reproduced·$3/$10 per M tokens
Consensus: LiteLLM · Cross-check: OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MMLU | 58.4accuracy (%) | unverified· optimizedT1 | 2024-02-26Stanford CRFM Leaderboard |
| MATH level 5 | 24.46accuracy (%) | reproduced· optimizedT1 | 2024-02-26Epoch AI |
| GPQA diamond | 18.35accuracy (%) | reproduced· optimizedT1 | 2024-02-26Epoch AI |
| OTIS Mock AIME 2024-2025 | 1.85accuracy (%) | reproduced· optimizedT1 | 2024-02-26Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by Mistral AI and released in 2024.
2 cited facts
On 4 tracked benchmarks a no-lead record is nearly uninformative — the sample is too small for the counts to carry signal; look at the individual scores instead. Because its results are independently reproduced, the scores are backed by outside evaluation, so they can be considered trustworthy.
3 cited facts
The model trails the leader by 68.0 points on average, a wide gap that places it well behind the frontier.
1 cited fact
5 facts cross-checked across data sources: 1 single-source, 4 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, openrouter_models
Mistral Large is an AI model developed by Mistral AI, released 2024-02-26. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Mistral Large has recorded scores on 4 benchmarks, each shown with its evidence status.
Mistral Large has recorded scores on 4 benchmarks — GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, MMLU. The full table above shows each score with its evidence status.
3 of 4 recorded scores (75%) are independently reproduced rather than self-reported by the lab.
Mistral Large has 4 tracked claims: 3 independently reproduced, 1 unverified.
Listed API pricing: $3 per million input tokens, $10 per million output tokens (prices disputed across sources). See the pricing block for the full breakdown.