01.AI·released 2023-11-021 source
Yi-34B benchmark scores: 5 benchmarks tracked. 40% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 5 benchmarks·best result 76 on GSM8K (unverified)·2 of 5 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| GSM8K | 76accuracy (%) | unverified· optimizedT1 | 2023-11-02Yi: Open Foundation Models by 01.AI |
| MMLU | 68.4accuracy (%) | unverified· optimizedT1 | 2023-11-02Stanford CRFM Leaderboard |
| BBH | 62.27accuracy (%) | unverified· optimizedT1 | 2023-11-02Yi: Open Foundation Models by 01.AI |
| MATH level 5 | 5.15accuracy (%) | reproduced· optimizedT1 | 2023-11-02Epoch AI |
| GPQA diamond | 0accuracy (%) | reproduced· optimizedT1 | 2023-11-02Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by 01.AI and released in 2023.
2 cited facts
Locating, not grading: 5 tracked benchmarks with no lead puts this model in the field, on a base too small to support anything sharper. Because the record is independently reproduced, outside evaluation backs up these scores, so they can be trusted.
3 cited facts
With an average gap of 48.17 points to the leader and no benchmark top scores, the model is well behind the leaders.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Yi-34B is an AI model developed by 01.AI, released 2023-11-02. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Yi-34B has recorded scores on 5 benchmarks, each shown with its evidence status.
Yi-34B has recorded scores on 5 benchmarks — GPQA diamond, MATH level 5, MMLU, BBH, GSM8K. The full table above shows each score with its evidence status.
2 of 5 recorded scores (40%) are independently reproduced rather than self-reported by the lab.
Yi-34B has 5 tracked claims: 2 independently reproduced, 3 unverified.