Zhipu AI·released 2023-06-241 source
chatglm2-6b benchmark scores: 7 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 7 benchmarks·best result 54.3 on LAMBADA (self-reported)·0 of 7 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| LAMBADA | 54.3accuracy (%) | self-reported· optimizedT1 | 2023-06-24Qwen Technical Report |
| ARC AI2 | 48accuracy (%) | self-reported· optimizedT1 | 2023-06-24Qwen Technical Report |
| HellaSwag | 42.67accuracy (%) | self-reported· optimizedT1 | 2023-06-24Qwen Technical Report |
| PIQA | 39.2accuracy (%) | self-reported· optimizedT1 | 2023-06-24Qwen Technical Report |
| GSM8K | 32.4accuracy (%) | unverified· optimizedT1 | 2023-06-24Baichuan 2: Open Large-scale Language Models |
| MMLU | 30.53accuracy (%) | self-reported· optimizedT1 | 2023-06-24Qwen Technical Report |
| BBH | 11.6accuracy (%) | unverified· optimizedT1 | 2023-06-24Baichuan 2: Open Large-scale Language Models |
Claims drawn from cited facts, not live model generation.
This model was developed by Zhipu AI and released in 2023.
2 cited facts
Its 7 tracked benchmarks make a moderate evidence base — broad enough to sketch a profile, thin enough that one strong or weak result still colors the whole record. Leading none of them puts it in the body of the field; that describes where it sits today, not what it can do under a different harness. The scores are self-reported, so they should be treated as vendor claims rather than independently verified results.
3 cited facts
The model trails the leader by an average of 49.91 points, a wide gap that places it well behind the front of the pack. It holds the top score on zero benchmarks, so it has no category leadership.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
chatglm2-6b is an AI model developed by Zhipu AI, released 2023-06-24. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
chatglm2-6b has recorded scores on 7 benchmarks, each shown with its evidence status.
chatglm2-6b has recorded scores on 7 benchmarks — PIQA, MMLU, ARC AI2, BBH, GSM8K, HellaSwag, and 1 more. The full table above shows each score with its evidence status.
0 of 7 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
chatglm2-6b has 7 tracked claims: 5 self-reported, 2 unverified.