Mistral AI·released 2026-04-291 source
Mistral Medium 3.5 benchmark scores: 5 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 5 benchmarks·best result 43.71 on WeirdML (unverified)·0 of 5 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| WeirdML | 43.71accuracy (%) | unverifiedT1 | 2026-04-29https://htihle.github.io/weirdml.html |
| Surface Evolver Bench | 26.88accuracy (%) | unverifiedT1 | 2026-04-29https://yhenon.github.io/surface-evolver-llm-eval/ |
| ProofBench | 9accuracy (%) | unverifiedT1 | 2026-04-29https://www.vals.ai/benchmarks/proof_bench |
| FrontierCode | 8accuracy (%) | unverifiedT1 | 2026-04-29https://cognition.com/frontiercode |
| CritPt | 0accuracy (%) | unverifiedT2 | 2026-04-29 |
Claims drawn from cited facts, not live model generation.
Developed by Mistral AI, this model was released in 2026; no parameter count is offered, so no scale-based efficiency or cost implications are stated.
2 cited facts
This model is tracked across five benchmarks. It currently holds no top score among them. Because the record is unverified, treat these scores as neither independently confirmed nor disputed by outside evaluation.
3 cited facts
This model trails the SOTA leader by an average of 56.83 points under the disclosed harness, a wide gap placing it well behind the front-runners. It holds the top score on none of the SOTA leaderboard benchmarks, confirming no category leadership on these tasks.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Mistral Medium 3.5 is an AI model developed by Mistral AI, released 2026-04-29. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Mistral Medium 3.5 has recorded scores on 5 benchmarks, each shown with its evidence status.
Mistral Medium 3.5 has recorded scores on 5 benchmarks — WeirdML, CritPt, FrontierCode, Surface Evolver Bench, ProofBench. The full table above shows each score with its evidence status.
0 of 5 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Mistral Medium 3.5 has 5 tracked claims: 5 unverified.