Qwen·released 2024-04-151 source
CodeQwen1.5-7B benchmark scores: 4 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 4 benchmarks·best result 37.7 on GSM8K (self-reported)·0 of 4 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| GSM8K | 37.7accuracy (%) | self-reported· optimizedT1 | 2024-04-15Qwen2.5-Coder Technical Report |
| MMLU | 20.67accuracy (%) | self-reported· optimizedT1 | 2024-04-15Qwen2.5-Coder Technical Report |
| Winogrande | 19.6accuracy (%) | self-reported· optimizedT1 | 2024-04-15Qwen2.5-Coder Technical Report |
| ARC AI2 | 14.27accuracy (%) | self-reported· optimizedT1 | 2024-04-15Qwen2.5-Coder Technical Report |
Claims drawn from cited facts, not live model generation.
This model was developed by Qwen and released in 2024.
2 cited facts
Sparse coverage — 4 benchmarks — locates the model in the data without characterizing it; hold conclusions until more scores land. None of those results is a current lead, a standings note that says little by itself about where the model's strengths sit. Because the scores are self-reported, they should be read as vendor claims rather than independently verified results.
3 cited facts
This model trails the state-of-the-art leader by an average of 64 points on its benchmarks, a wide gap that places it well behind the leading systems. It holds the top score on zero benchmarks, indicating it is not a category leader.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
CodeQwen1.5-7B is an AI model developed by Qwen, released 2024-04-15. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
CodeQwen1.5-7B has recorded scores on 4 benchmarks, each shown with its evidence status.
CodeQwen1.5-7B has recorded scores on 4 benchmarks — Winogrande, MMLU, ARC AI2, GSM8K. The full table above shows each score with its evidence status.
0 of 4 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
CodeQwen1.5-7B has 4 tracked claims: 4 self-reported.