MosaicML·released 2023-05-051 source
MPT-7B benchmark scores: 10 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 10 benchmarks·best result 70 on LAMBADA (self-reported)·0 of 10 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| LAMBADA | 70accuracy (%) | self-reported· optimizedT1 | 2023-05-05Qwen Technical Report |
| HellaSwag | 68.53accuracy (%) | self-reported· optimizedT1 | 2023-05-05Qwen Technical Report |
| TriviaQA | 61.6accuracy (%) | unverified· optimizedT1 | 2023-05-05Llama 2: Open Foundation and Fine-Tuned Chat Models |
| PIQA | 61.2accuracy (%) | self-reported· optimizedT1 | 2023-05-05Qwen Technical Report |
| Winogrande | 37.2accuracy (%) | unverified· optimizedT1 | 2023-05-05Llama 2: Open Foundation and Fine-Tuned Chat Models |
| OpenBookQA | 35.2accuracy (%) | unverified· optimizedT1 | 2023-05-05Llama 2: Open Foundation and Fine-Tuned Chat Models |
| ARC AI2 | 23.47accuracy (%) | self-reported· optimizedT1 | 2023-05-05Qwen Technical Report |
| BBH | 14.13accuracy (%) | unverified· optimizedT1 | 2023-05-05Baichuan 2: Open Large-scale Language Models |
| GSM8K | 9.1accuracy (%) | unverified· optimizedT1 | 2023-05-05Baichuan 2: Open Large-scale Language Models |
| MMLU | 7.73accuracy (%) | self-reported· optimizedT1 | 2023-05-05Qwen Technical Report |
Claims drawn from cited facts, not live model generation.
This model was developed by MosaicML and released in 2023.
2 cited facts
Reasonable breadth at 10 tracked scores favors reading trends across the row over fixating on any cell. None of them shows this model on top right now — a statement with a short shelf life, rewritten whenever the tables update. The record is self-reported, meaning the numbers are vendor-claimed and not yet independently confirmed, so they should be read as claims rather than verified results.
3 cited facts
With an average score gap of 46.99 points behind the leader, this model falls into the 'well behind the leaders' band.
1 cited fact
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
MPT-7B is an AI model developed by MosaicML, released 2023-05-05. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
MPT-7B has recorded scores on 10 benchmarks, each shown with its evidence status.
MPT-7B has recorded scores on 10 benchmarks — PIQA, OpenBookQA, Winogrande, TriviaQA, MMLU, ARC AI2, and 4 more. The full table above shows each score with its evidence status.
0 of 10 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
MPT-7B has 10 tracked claims: 5 self-reported, 5 unverified.