Mistral AI·released 2024-05-271 source
Mistral 7B v0.3 benchmark scores: 6 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 6 benchmarks·best result 46.53 on MMLU (unverified)·0 of 6 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MMLU | 46.53accuracy (%) | unverifiedT1 | 2024-05-27Stanford CRFM Leaderboard |
| DTBench | 4.23accuracy (%) | unverifiedT1 | 2023-09-27https://conceptualreasoning.ai/dtbench |
| MATH level 5 | 3.68accuracy (%) | unverifiedT2 | 2024-05-27 |
| OTIS Mock AIME 2024-2025 | 0.18accuracy (%) | unverifiedT2 | 2024-05-27 |
| GPQA diamond | 0accuracy (%) | unverifiedT2 | 2024-05-27 |
| Chess Puzzles | 0accuracy (%) | unverifiedT2 | 2024-05-27 |
Claims drawn from cited facts, not live model generation.
Developed by Mistral AI, this model was released in 2024.
2 cited facts
This model is evaluated on 6 tracked benchmarks, but it currently holds no top score on any of them. Because its record is unverified, no outside evaluation has either confirmed or disputed the reported numbers, so they should be read as claims rather than settled results.
3 cited facts
On average this model trails the SOTA leader by 81.64 points under the disclosed harness — a wide gap that places it well behind the front of the pack rather than at or near the frontier. Consistent with that, it holds the top score on 0 benchmarks, meaning it currently shows no category leadership on any of its evaluated tasks.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Mistral 7B v0.3 is an AI model developed by Mistral AI, released 2024-05-27. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Mistral 7B v0.3 has recorded scores on 6 benchmarks, each shown with its evidence status.
Mistral 7B v0.3 has recorded scores on 6 benchmarks — GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, DTBench, MMLU, Chess Puzzles. The full table above shows each score with its evidence status.
0 of 6 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Mistral 7B v0.3 has 6 tracked claims: 6 unverified.