DeepSeek·released 2025-01-201 source
DeepSeek-R1-Distill-Llama-70B benchmark scores: 5 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 5 benchmarks·best result 1633 on Codeforces rating (self-reported)·0 of 5 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| Codeforces rating | 1633rating | self-reported· optimizedT1 | 2025-01-22DeepSeek AI |
| MATH-500 | 94.5pass@1 (%) | self-reported· optimizedT1 | 2025-01-22DeepSeek AI |
| AIME | 70pass@1 (%) | self-reported· optimizedT1 | 2025-01-22DeepSeek AI |
| GPQA diamond | 65.2accuracy (%) | self-reported· optimizedT1 | 2025-01-22DeepSeek AI |
| LiveCodeBench | 57.5pass@1 (%) | self-reported· optimizedT1 | 2025-01-22DeepSeek AI |
Claims drawn from cited facts, not live model generation.
Developed by DeepSeek, this model operates at a large scale of 70.55 billion parameters, a configuration that trades efficiency for more headroom and expense at the large end. It was introduced in 2025.
3 cited facts
Evaluated across 5 benchmarks, the model holds no current top score. Because the record is self-reported, these figures should be treated as vendor claims rather than independently confirmed results.
3 cited facts
The model trails the current leader by an average of 88.92 points across its benchmarks, as indicated by avg sota gap. This wide margin places it well behind the leaders.
1 cited fact
6 facts cross-checked across data sources: 6 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: curated_benchmark_scores, huggingface_models, openrouter_models
DeepSeek-R1-Distill-Llama-70B is an AI model developed by DeepSeek, released 2025-01-20. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
DeepSeek-R1-Distill-Llama-70B has recorded scores on 5 benchmarks, each shown with its evidence status.
DeepSeek-R1-Distill-Llama-70B has recorded scores on 5 benchmarks — AIME, Codeforces rating, GPQA diamond, LiveCodeBench, MATH-500. The full table above shows each score with its evidence status.
0 of 5 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
DeepSeek-R1-Distill-Llama-70B has 5 tracked claims: 5 self-reported.