Mistral AI·released 2024-07-181 source
Mistral NeMo benchmark scores: 5 benchmarks tracked. 40% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 5 benchmarks·best result 84.2 on GSM8K (self-reported)·2 of 5 independently reproduced·$0.15/$0.15 per M tokens
Source: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| GSM8K | 84.2accuracy (%) | self-reported· optimizedT1 | 2024-07-18Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| PIQA | 67accuracy (%) | self-reported· optimizedT1 | 2024-07-18Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| Balrog | 17.6accuracy (%) | unverifiedT1 | 2024-07-18Balrog Leaderboard |
| MATH level 5 | 10.83accuracy (%) | reproduced· optimizedT1 | 2024-07-18Epoch AI |
| GPQA diamond | 6.52accuracy (%) | reproduced· optimizedT1 | 2024-07-18Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by Mistral AI and released in 2024.
2 cited facts
Just 5 tracked scores back this record — small enough that a single result still swings the picture. Holding none is a snapshot of today's table tops — it neither places the model in the field nor forecasts the next update. Because its scores are independently reproduced, the record can be trusted as verified.
3 cited facts
This model trails the state‑of‑the‑art by an average of 46.8 points and holds zero top scores, placing it well behind the leaders.
2 cited facts
5 facts cross-checked across data sources: 1 corroborated, 1 single-source, 3 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, modelsdev_models, openrouter_models
Mistral NeMo is an AI model developed by Mistral AI, released 2024-07-18. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Mistral NeMo has recorded scores on 5 benchmarks, each shown with its evidence status.
Mistral NeMo has recorded scores on 5 benchmarks — PIQA, GPQA diamond, MATH level 5, Balrog, GSM8K. The full table above shows each score with its evidence status.
2 of 5 recorded scores (40%) are independently reproduced rather than self-reported by the lab.
Mistral NeMo has 5 tracked claims: 2 independently reproduced, 2 self-reported, 1 unverified.
Listed API pricing: $0.15 per million input tokens, $0.15 per million output tokens (prices disputed across sources). See the pricing block for the full breakdown.