DeepSeek·released 2025-01-201 source
DeepSeek-R1-Distill-Llama-8B benchmark scores: 5 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 5 benchmarks·best result 1205 on Codeforces rating (self-reported)·0 of 5 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| Codeforces rating | 1205rating | self-reported· optimizedT1 | 2025-01-22DeepSeek AI |
| MATH-500 | 89.1pass@1 (%) | self-reported· optimizedT1 | 2025-01-22DeepSeek AI |
| AIME | 50.4pass@1 (%) | self-reported· optimizedT1 | 2025-01-22DeepSeek AI |
| GPQA diamond | 49accuracy (%) | self-reported· optimizedT1 | 2025-01-22DeepSeek AI |
| LiveCodeBench | 39.6pass@1 (%) | self-reported· optimizedT1 | 2025-01-22DeepSeek AI |
Claims drawn from cited facts, not live model generation.
This model was developed by DeepSeek and released in 2025.
2 cited facts
The model has been evaluated across five tracked benchmarks, yet it currently holds no top score among them. These results are self-reported, meaning the numbers are vendor-claimed and not yet independently confirmed, so they should be read as claims rather than verified results.
3 cited facts
The model trails the current leader by an average of 186.34 points, a wide gap that places it well behind the frontier. It also holds the top score on zero benchmarks, indicating it has not yet achieved category leadership on these tasks.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: curated_benchmark_scores, curated_models
DeepSeek-R1-Distill-Llama-8B is an AI model developed by DeepSeek, released 2025-01-20. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
DeepSeek-R1-Distill-Llama-8B has recorded scores on 5 benchmarks, each shown with its evidence status.
DeepSeek-R1-Distill-Llama-8B has recorded scores on 5 benchmarks — AIME, Codeforces rating, GPQA diamond, LiveCodeBench, MATH-500. The full table above shows each score with its evidence status.
0 of 5 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
DeepSeek-R1-Distill-Llama-8B has 5 tracked claims: 5 self-reported.