DeepSeek·released 2025-01-201 source
DeepSeek-R1-Distill-Qwen-32B benchmark scores: 5 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 5 benchmarks·best result 1691 on Codeforces rating (self-reported)·0 of 5 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| Codeforces rating | 1691rating | self-reported· optimizedT1 | 2025-01-22DeepSeek AI |
| MATH-500 | 94.3pass@1 (%) | self-reported· optimizedT1 | 2025-01-22DeepSeek AI |
| AIME | 72.6pass@1 (%) | self-reported· optimizedT1 | 2025-01-22DeepSeek AI |
| GPQA diamond | 62.1accuracy (%) | self-reported· optimizedT1 | 2025-01-22DeepSeek AI |
| LiveCodeBench | 57.2pass@1 (%) | self-reported· optimizedT1 | 2025-01-22DeepSeek AI |
Claims drawn from cited facts, not live model generation.
Developed by DeepSeek, this model is a large model released in 2025 with 32.76 billion parameters; that scale affords substantial headroom for complex tasks, but also entails higher computational expense.
3 cited facts
This model is tracked across five benchmarks, but it holds no current top score on any of them. Because these figures are self-reported, they should be read as claims rather than independently verified results.
3 cited facts
On the disclosed harness, the model trails the leader by an average of 77.58 points, a wide gap that places it well behind the state of the art. It holds the top score on none of the benchmarks, so there is no category leadership yet.
2 cited facts
2 facts cross-checked across data sources: 2 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: curated_benchmark_scores, curated_capability_claims, huggingface_models
DeepSeek-R1-Distill-Qwen-32B is an AI model developed by DeepSeek, released 2025-01-20. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
DeepSeek-R1-Distill-Qwen-32B has recorded scores on 5 benchmarks, each shown with its evidence status.
DeepSeek-R1-Distill-Qwen-32B has recorded scores on 5 benchmarks — AIME, Codeforces rating, GPQA diamond, LiveCodeBench, MATH-500. The full table above shows each score with its evidence status.
0 of 5 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
DeepSeek-R1-Distill-Qwen-32B has 5 tracked claims: 5 self-reported.