Verdict
DeepSeek-R1-Distill-Llama-70B leads 3–2 across 5 shared benchmarks.
DeepSeek-R1-Distill-Llama-70B 3 · DeepSeek-R1-Distill-Qwen-32B 2 · higher isn't always better — check the caveats below.
Reported scores — protocols may differ. How we compare
Coding: Evenly split 1–1 across 2 shared coding benchmarks (largest gap: Codeforces rating, 1691 vs 1633 — single evaluator).
Vendor claims — compare against the measured scores above. A claim is what a developer says about its own model, not an independent measurement.
No vendor claims on file for DeepSeek-R1-Distill-Llama-70B.
Reasoning
“Using Qwen2.5-32B as the base model, direct distillation from DeepSeek-R1 outperforms applying RL on it.”
DeepSeek-R1-Distill-Llama-70B leads 3–2 across 5 shared benchmarks. DeepSeek-R1-Distill-Llama-70B leads 3 benchmarks and DeepSeek-R1-Distill-Qwen-32B leads 2. Higher isn't always better — see the integrity caveats on each benchmark.
DeepSeek-R1-Distill-Llama-70B and DeepSeek-R1-Distill-Qwen-32B have 5 benchmarks in common in our data — those are the only rows where a direct, apples-to-apples comparison is drawn.