Mistral AI·released 2023-12-111 source
Mixtral 8x7B benchmark scores: 11 benchmarks tracked. 18% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 11 benchmarks·best result 83.07 on ARC AI2 (unverified)·2 of 11 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| ARC AI2 | 83.07accuracy (%) | unverified· optimizedT1 | 2023-12-11Mixtral of Experts |
| HellaSwag | 82.27accuracy (%) | unverified· optimizedT1 | 2023-12-11Mixtral of Experts |
| TriviaQA | 82.2accuracy (%) | self-reported· optimizedT1 | 2023-12-11Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| OpenBookQA | 81.07accuracy (%) | self-reported· optimizedT1 | 2023-12-11Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| GSM8K | 74.4accuracy (%) | unverified· optimizedT1 | 2023-12-11Mixtral of Experts |
| PIQA | 67.2accuracy (%) | unverified· optimizedT1 | 2023-12-11Mixtral of Experts |
| MMLU | 60.8accuracy (%) | unverified· optimizedT1 | 2023-12-11Mixtral of Experts |
| Winogrande | 54.4accuracy (%) | unverified· optimizedT1 | 2023-12-11Mixtral of Experts |
| ANLI | 32.8accuracy (%) | self-reported· optimizedT1 | 2023-12-11Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| MATH level 5 | 9.95accuracy (%) | reproduced· optimizedT1 | 2023-12-11Epoch AI |
| GPQA diamond | 7.45accuracy (%) | reproduced· optimizedT1 | 2023-12-11Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by Mistral AI and released in 2023.
2 cited facts
This model is scored on 11 tracked benchmarks and holds no current top score; because its scores have been independently reproduced, the record is trustworthy.
3 cited facts
The model trails the state-of-the-art by 25.92 points on average, a wide gap indicating it is well behind the leaders, and it holds the top score on none of the benchmarks, meaning it has not achieved category leadership.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: curated_capability_claims, epoch_benchmarks
Mixtral 8x7B is an AI model developed by Mistral AI, released 2023-12-11. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Mixtral 8x7B has recorded scores on 11 benchmarks, each shown with its evidence status.
Mixtral 8x7B has recorded scores on 11 benchmarks — PIQA, OpenBookQA, Winogrande, TriviaQA, GPQA diamond, MATH level 5, and 5 more. The full table above shows each score with its evidence status.
2 of 11 recorded scores (18%) are independently reproduced rather than self-reported by the lab.
Mixtral 8x7B has 11 tracked claims: 2 independently reproduced, 3 self-reported, 6 unverified.