01.AI·released 2023-11-021 source
Yi 6B benchmark scores: 6 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 6 benchmarks·best result 65.87 on HellaSwag (unverified)·0 of 6 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| HellaSwag | 65.87accuracy (%) | unverified· optimizedT1 | 2023-11-02Yi: Open Foundation Models by 01.AI |
| MMLU | 52accuracy (%) | unverified· optimizedT1 | 2023-11-02Stanford CRFM Leaderboard |
| GSM8K | 44.9accuracy (%) | unverified· optimizedT1 | 2023-11-02Yi: Open Foundation Models by 01.AI |
| Winogrande | 42.6accuracy (%) | unverified· optimizedT1 | 2023-11-02Yi: Open Foundation Models by 01.AI |
| ARC AI2 | 33.73accuracy (%) | unverified· optimizedT1 | 2023-11-02Yi: Open Foundation Models by 01.AI |
| BBH | 29.6accuracy (%) | unverified· optimizedT1 | 2023-11-02Yi: Open Foundation Models by 01.AI |
Claims drawn from cited facts, not live model generation.
This model was developed by 01.AI and released in 2023.
2 cited facts
Benchmark coverage sits at 6 here — enough to sketch behavior, well short of settling it. Standing is all this fixes: the model leads none of its measured boards, which says where it sits today and nothing further. Because the record is unverified, the scores are neither confirmed nor disputed by outside evaluation and should be treated as unverified claims.
3 cited facts
With an average gap of 43.15 points to the leader and no benchmarks where it holds the top score, this model is well behind the frontier.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Yi 6B is an AI model developed by 01.AI, released 2023-11-02. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Yi 6B has recorded scores on 6 benchmarks, each shown with its evidence status.
Yi 6B has recorded scores on 6 benchmarks — Winogrande, MMLU, ARC AI2, BBH, GSM8K, HellaSwag. The full table above shows each score with its evidence status.
0 of 6 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Yi 6B has 6 tracked claims: 6 unverified.