Qwen·released 2023-11-301 source
Qwen-1_8B benchmark scores: 6 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 6 benchmarks·best result 58.4 on LAMBADA (self-reported)·0 of 6 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| LAMBADA | 58.4accuracy (%) | self-reported· optimizedT1 | 2023-11-30Qwen Technical Report |
| PIQA | 46.6accuracy (%) | self-reported· optimizedT1 | 2023-11-30Qwen Technical Report |
| ARC AI2 | 37.6accuracy (%) | self-reported· optimizedT1 | 2023-11-30Qwen Technical Report |
| GSM8K | 21.2accuracy (%) | self-reported· optimizedT1 | 2023-11-30Qwen Technical Report |
| MMLU | 4.27accuracy (%) | self-reported· optimizedT1 | 2023-11-30Qwen Technical Report |
| BBH | 4.27accuracy (%) | self-reported· optimizedT1 | 2023-11-30Qwen Technical Report |
Claims drawn from cited facts, not live model generation.
This model was developed by Qwen and released in 2023.
2 cited facts
A field position on a modest base: scored across 6 tracked benchmarks and topping none, this model is located without being capped. All scores are self-reported, meaning they are vendor-claimed and not independently confirmed, so they should be read as claims rather than verified results.
3 cited facts
This model trails the leader by an average of 57.0 points across benchmarks, placing it well behind the frontier. It holds the top score on none of the benchmarks, indicating it is not a category leader.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Qwen-1_8B is an AI model developed by Qwen, released 2023-11-30. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Qwen-1_8B has recorded scores on 6 benchmarks, each shown with its evidence status.
Qwen-1_8B has recorded scores on 6 benchmarks — PIQA, MMLU, ARC AI2, BBH, GSM8K, LAMBADA. The full table above shows each score with its evidence status.
0 of 6 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Qwen-1_8B has 6 tracked claims: 6 self-reported.