Qwen·released 2024-09-181 source
Qwen2.5-Coder (1.5B) benchmark scores: 5 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 5 benchmarks·best result 65.8 on GSM8K (self-reported)·0 of 5 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| GSM8K | 65.8accuracy (%) | self-reported· optimizedT1 | 2024-09-18Qwen2.5-Coder Technical Report |
| HellaSwag | 49.07accuracy (%) | self-reported· optimizedT1 | 2024-09-18Qwen2.5-Coder Technical Report |
| MMLU | 38.13accuracy (%) | self-reported· optimizedT1 | 2024-09-18Qwen2.5-Coder Technical Report |
| ARC AI2 | 26.93accuracy (%) | self-reported· optimizedT1 | 2024-09-18Qwen2.5-Coder Technical Report |
| Winogrande | 21.4accuracy (%) | self-reported· optimizedT1 | 2024-09-18Qwen2.5-Coder Technical Report |
Claims drawn from cited facts, not live model generation.
This model was developed by Qwen. It was released in 2024.
2 cited facts
This model is scored on 5 tracked benchmarks but holds no current top score. Because the scores are self-reported, they should be treated as claims rather than independently verified results.
3 cited facts
This model trails the state-of-the-art leader by an average of 48.13 points across benchmarks, a wide gap that places it well behind the frontrunners, and it holds the top score on none of the benchmarks, indicating no category leadership.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Qwen2.5-Coder (1.5B) is an AI model developed by Qwen, released 2024-09-18. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Qwen2.5-Coder (1.5B) has recorded scores on 5 benchmarks, each shown with its evidence status.
Qwen2.5-Coder (1.5B) has recorded scores on 5 benchmarks — Winogrande, MMLU, ARC AI2, GSM8K, HellaSwag. The full table above shows each score with its evidence status.
0 of 5 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Qwen2.5-Coder (1.5B) has 5 tracked claims: 5 self-reported.