Baichuan AI·released 2023-09-061 source
Baichuan2-13B benchmark scores: 7 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 7 benchmarks·best result 74 on LAMBADA (self-reported)·0 of 7 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| LAMBADA | 74accuracy (%) | self-reported· optimizedT1 | 2023-09-06Qwen Technical Report |
| HellaSwag | 61.07accuracy (%) | self-reported· optimizedT1 | 2023-09-06Qwen Technical Report |
| PIQA | 56.2accuracy (%) | self-reported· optimizedT1 | 2023-09-06Qwen Technical Report |
| GSM8K | 52.8accuracy (%) | self-reported· optimizedT1 | 2023-09-06Nemotron-4 15B Technical Report |
| MMLU | 45.6accuracy (%) | unverified· optimizedT1 | 2023-09-06Baichuan 2: Open Large-scale Language Models |
| BBH | 32accuracy (%) | self-reported· optimizedT1 | 2023-09-06Nemotron-4 15B Technical Report |
| ARC AI2 | 17.33accuracy (%) | self-reported· optimizedT1 | 2023-09-06Qwen Technical Report |
Claims drawn from cited facts, not live model generation.
This model was developed by Baichuan AI and released in 2023.
2 cited facts
Coverage across 7 benchmarks gives a mid-sized record; topping none of them places it in the pack, which is where most of any tracked field sits at a given moment. Because these results are self-reported and not independently verified, they should be treated as claims rather than confirmed achievements.
3 cited facts
The model trails the state-of-the-art by an average gap of 38.44 points, placing it well behind the leaders. No benchmark in its category puts it on top. That is a position in the field, not a reading on the model — this count records whether it leads, never how far behind it runs.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Baichuan2-13B is an AI model developed by Baichuan AI, released 2023-09-06. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Baichuan2-13B has recorded scores on 7 benchmarks, each shown with its evidence status.
Baichuan2-13B has recorded scores on 7 benchmarks — PIQA, MMLU, ARC AI2, BBH, GSM8K, HellaSwag, and 1 more. The full table above shows each score with its evidence status.
0 of 7 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Baichuan2-13B has 7 tracked claims: 6 self-reported, 1 unverified.