MosaicML·released 2023-06-221 source
MPT-30B benchmark scores: 8 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 8 benchmarks·best result 73.6 on TriviaQA (unverified)·0 of 8 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| TriviaQA | 73.6accuracy (%) | unverified· optimizedT1 | 2023-06-22Llama 2: Open Foundation and Fine-Tuned Chat Models |
| PIQA | 63.8accuracy (%) | unverified· optimizedT1 | 2023-06-22Llama 2: Open Foundation and Fine-Tuned Chat Models |
| Winogrande | 42accuracy (%) | unverified· optimizedT1 | 2023-06-22Llama 2: Open Foundation and Fine-Tuned Chat Models |
| OpenBookQA | 36accuracy (%) | unverified· optimizedT1 | 2023-06-22Llama 2: Open Foundation and Fine-Tuned Chat Models |
| GSM8K | 34.4accuracy (%) | unverified· optimizedT1 | 2023-06-22Stanford HELM |
| ARC AI2 | 34.13accuracy (%) | unverified· optimizedT1 | 2023-06-22Llama 2: Open Foundation and Fine-Tuned Chat Models |
| MMLU | 30.53accuracy (%) | self-reported· optimizedT1 | 2023-06-22Qwen Technical Report |
| BBH | 17.33accuracy (%) | self-reported· optimizedT1 | 2023-06-22Qwen Technical Report |
Claims drawn from cited facts, not live model generation.
This model was developed by MosaicML and released in 2023.
2 cited facts
This model is scored on 8 tracked benchmarks. Absence of a lead excludes exactly this much: current leadership. Everything else about the model's standing stays open. The scores are self-reported, so treat them as claims rather than verified results.
3 cited facts
With an average gap of 44.1 points behind the state-of-the-art leader, the model is well behind the front-of-pack. It holds the top score on zero benchmarks, indicating it has not achieved category leadership.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
MPT-30B is an AI model developed by MosaicML, released 2023-06-22. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
MPT-30B has recorded scores on 8 benchmarks, each shown with its evidence status.
MPT-30B has recorded scores on 8 benchmarks — PIQA, OpenBookQA, Winogrande, TriviaQA, MMLU, ARC AI2, and 2 more. The full table above shows each score with its evidence status.
0 of 8 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
MPT-30B has 8 tracked claims: 2 self-reported, 6 unverified.