01.AI·released 2024-03-011 source
Yi-9B benchmark scores: 4 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 4 benchmarks·best result 68.53 on HellaSwag (unverified)·0 of 4 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| HellaSwag | 68.53accuracy (%) | unverified· optimizedT1 | 2024-03-01Yi: Open Foundation Models by 01.AI |
| MMLU | 57.87accuracy (%) | unverified· optimizedT1 | 2024-03-01Yi: Open Foundation Models by 01.AI |
| Winogrande | 46accuracy (%) | unverified· optimizedT1 | 2024-03-01Yi: Open Foundation Models by 01.AI |
| ARC AI2 | 40.8accuracy (%) | unverified· optimizedT1 | 2024-03-01Yi: Open Foundation Models by 01.AI |
Claims drawn from cited facts, not live model generation.
This model was developed by 01.AI and released in 2024.
2 cited facts
Too small a sample to conclude much: 4 benchmarks with no lead is a thin placement that more measurement could revise in either direction. The performance record is unverified, meaning it has neither been confirmed nor disputed by outside evaluation.
3 cited facts
The model trails the state-of-the-art by an average of 34.2 points on its benchmarks, placing it well behind the frontier group. It does not achieve the top score on any of the benchmarks, confirming that it is not a category leader.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Yi-9B is an AI model developed by 01.AI, released 2024-03-01. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Yi-9B has recorded scores on 4 benchmarks, each shown with its evidence status.
Yi-9B has recorded scores on 4 benchmarks — Winogrande, MMLU, ARC AI2, HellaSwag. The full table above shows each score with its evidence status.
0 of 4 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Yi-9B has 4 tracked claims: 4 unverified.