DeepSeek·released 2025-01-201 source
DeepSeek-R1-Distill-Qwen-32B benchmark scores: 9 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 8 benchmarks·best result 1691 on Codeforces rating (self-reported)·0 of 9 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| Codeforces rating | 1691rating | self-reported· optimizedT1 | 2025-01-22DeepSeek AI |
| MATH-500 | 94.3pass@1 (%) | self-reported· optimizedT1 | 2025-01-22DeepSeek AI |
| AIME | 72.6pass@1 (%) | self-reported· optimizedT1 | 2025-01-22DeepSeek AI |
| GPQA diamond | 62.1accuracy (%) | self-reported· optimizedT1 | 2025-01-22DeepSeek AI |
| LiveCodeBench | 57.2pass@1 (%) | self-reported· optimizedT1 | 2025-01-22DeepSeek AI |
| OTIS Mock AIME 2024-2025 | 55.51accuracy (%) | unverifiedT2 | 2025-01-20 |
| GPQA diamond | 52.19accuracy (%) | unverifiedT2 | 2025-01-20 |
| Balrog | 19.5accuracy (%) | unverifiedT1 | 2025-01-20https://balrogai.com/ |
| Chess Puzzles | 0accuracy (%) | unverifiedT2 | 2025-01-20 |
Claims drawn from cited facts, not live model generation.
This model was developed by DeepSeek and released in 2025.
2 cited facts
This model is scored on 8 tracked benchmarks, and it currently holds no top score on any of them. Its record is self-reported: the figures are vendor claims that have not been independently confirmed, so they should be read as claims rather than verified results.
3 cited facts
On average it trails the leader by 67.85 points on its benchmarks, a wide gap that reads as well behind the front of the pack. It holds the top score on none of its benchmarks, meaning it demonstrates no genuine category leadership yet.
2 cited facts
Benchmarks with more than one recorded measurement for this model — every one shown, not just the headline number.
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: curated_benchmark_scores, curated_capability_claims, epoch_benchmarks
DeepSeek-R1-Distill-Qwen-32B is an AI model developed by DeepSeek, released 2025-01-20. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
DeepSeek-R1-Distill-Qwen-32B has recorded scores on 8 benchmarks, each shown with its evidence status.
DeepSeek-R1-Distill-Qwen-32B has recorded scores on 9 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, Chess Puzzles, Balrog, AIME, Codeforces rating, and 3 more. The full table above shows each score with its evidence status.
0 of 9 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
DeepSeek-R1-Distill-Qwen-32B has 8 tracked claims: 5 self-reported, 3 unverified.