Qwen·released 2023-09-281 source
Qwen-7B benchmark scores: 6 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 6 benchmarks·best result 67.9 on LAMBADA (self-reported)·0 of 6 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| LAMBADA | 67.9accuracy (%) | self-reported· optimizedT1 | 2023-09-28Qwen Technical Report |
| ARC AI2 | 67.07accuracy (%) | self-reported· optimizedT1 | 2023-09-28Qwen Technical Report |
| PIQA | 55.8accuracy (%) | self-reported· optimizedT1 | 2023-09-28Qwen Technical Report |
| GSM8K | 51.7accuracy (%) | unverified· optimizedT1 | 2023-09-28Yi: Open Foundation Models by 01.AI |
| MMLU | 26.67accuracy (%) | self-reported· optimizedT1 | 2023-09-28Qwen Technical Report |
| BBH | 26.67accuracy (%) | self-reported· optimizedT1 | 2023-09-28Qwen Technical Report |
Claims drawn from cited facts, not live model generation.
This model was developed by Qwen. It was released in 2023.
2 cited facts
Depth of testing here is 6 tracked benchmarks — a partial view, so favor consistency across the set over any standalone result. No top result sits in that set for this model; treat the fact as a snapshot, since leadership counts move as new scores land. Because its results are self-reported, they should be regarded as vendor claims rather than independently verified numbers.
3 cited facts
This model trails the state-of-the-art by 36.43 points on average, a wide gap that places it well behind the leaders. It holds the top score on none of the benchmarks, indicating it is not a category leader.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Qwen-7B is an AI model developed by Qwen, released 2023-09-28. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Qwen-7B has recorded scores on 6 benchmarks, each shown with its evidence status.
Qwen-7B has recorded scores on 6 benchmarks — PIQA, MMLU, ARC AI2, BBH, GSM8K, LAMBADA. The full table above shows each score with its evidence status.
0 of 6 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Qwen-7B has 6 tracked claims: 5 self-reported, 1 unverified.