DeepSeek·released 2024-01-251 source
DeepSeek Coder 6.7B benchmark scores: 4 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 4 benchmarks·best result 21.3 on GSM8K (self-reported)·0 of 4 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| GSM8K | 21.3accuracy (%) | self-reported· optimizedT1 | 2024-01-25Qwen2.5-Coder Technical Report |
| Winogrande | 15.2accuracy (%) | self-reported· optimizedT1 | 2024-01-25Qwen2.5-Coder Technical Report |
| MMLU | 15.2accuracy (%) | self-reported· optimizedT1 | 2024-01-25Qwen2.5-Coder Technical Report |
| ARC AI2 | 15.2accuracy (%) | self-reported· optimizedT1 | 2024-01-25Qwen2.5-Coder Technical Report |
Claims drawn from cited facts, not live model generation.
This model was developed by DeepSeek and released in 2024.
2 cited facts
Tracking across 4 benchmarks is a starting footprint rather than a full evaluation, so weight this record lightly. Holding no top score is the default state for a tracked model — leadership is scarce by construction — so read this as unexceptional rather than damning. Because these figures are self-reported, they should be treated as claims rather than independently confirmed results.
3 cited facts
The model trails the state-of-the-art by an average of 70.34 points, a wide gap indicating it is well behind the leaders. It holds the top score on none of the benchmarks, so it does not achieve genuine category leadership.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
DeepSeek Coder 6.7B is an AI model developed by DeepSeek, released 2024-01-25. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
DeepSeek Coder 6.7B has recorded scores on 4 benchmarks, each shown with its evidence status.
DeepSeek Coder 6.7B has recorded scores on 4 benchmarks — Winogrande, MMLU, ARC AI2, GSM8K. The full table above shows each score with its evidence status.
0 of 4 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
DeepSeek Coder 6.7B has 4 tracked claims: 4 self-reported.