DeepSeek·released 2025-01-20✓ 2 sources
DeepSeek-R1 benchmark scores: 18 benchmarks tracked, leading 4. 17% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 4 of 17 benchmarks·best result 2029 on Codeforces rating (self-reported)·3 of 18 independently reproduced·$0.28/$0.42 per M tokens
Head-to-headDeepSeek-R1 vs DeepSeek-R1-Distill-Qwen-32BConsensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
Leads on: AIME, Codeforces rating, LiveCodeBench, MATH-500
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| Codeforces rating | 2029rating | self-reported· optimizedT1 | 2025-01-22DeepSeek AI |
| MATH-500 | 97.3pass@1 (%) | self-reported· optimizedT1 | 2025-01-22DeepSeek AI |
| MATH level 5 | 93.05accuracy (%) | reproduced· optimizedT1 | 2025-01-20Epoch AI |
| Lech Mazur Writing | 83accuracy (%) | unverifiedT1 | 2025-01-20lechmazur/writing Github repository |
| AIME | 79.8pass@1 (%) | self-reported· optimizedT1 | 2025-01-22DeepSeek AI |
| GPQA diamond | 71.5accuracy (%) | self-reported· optimizedT1 | 2025-01-22DeepSeek AI |
| Fiction.LiveBench | 69.4accuracy (%) | unverifiedT1 | 2025-01-20Fiction.live leaderboard |
| LiveCodeBench | 65.9pass@1 (%) | self-reported· optimizedT1 | 2025-01-22DeepSeek AI |
| GPQA diamond | 62.29accuracy (%) | reproduced· optimizedT1 | 2025-01-20Epoch AI |
| Aider polyglot | 56.9accuracy (%) | unverified· optimizedT1 | 2025-01-20Aider LLM Leaderboards |
| OTIS Mock AIME 2024-2025 | 53.29accuracy (%) | reproduced· optimizedT1 | 2025-01-20Epoch AI |
| METR Time Horizons | 51.93accuracy (%) | unverifiedT1 | 2025-01-20METR - Measuring AI Ability to Complete Long Tasks |
| WeirdML | 36.49accuracy (%) | unverifiedT1 | 2025-01-20WeirdML Leaderboard |
| Balrog | 34.9accuracy (%) | unverifiedT1 | 2025-01-20Balrog Leaderboard |
| SimpleBench | 17.08accuracy (%) | unverifiedT1 | 2025-01-20SimpleBench Leaderboard |
| ARC-AGI | 15.8accuracy (%) | unverified· optimizedT1 | 2025-01-20ARC Prize Leaderboard |
| ARC-AGI-2 | 1.3accuracy (%) | unverifiedT2 | 2025-01-20 |
| CritPt | 1.1accuracy (%) | unverifiedT2 | 2025-01-20 |
Claims drawn from cited facts, not live model generation.
Developed by DeepSeek and released in 2025, this model is frontier-scale. With 684.53 billion parameters, this model's frontier scale implies significant headroom but also high computational expense, a trade-off inherent to large models.
4 cited facts
This model is tracked across 17 benchmarks, and it currently holds the top score on 4 of them. However, the record is contradicted by outside evaluation, so its self-reported numbers should be treated with caution rather than as verified results.
3 cited facts
This model trails its benchmark leaders by an average of 29.83 points, a wide gap that places it well behind the front of the pack rather than at the frontier. Although it reports the top score on 4 benchmarks, that does not translate into genuine category leadership given the substantial average shortfall.
2 cited facts
Benchmarks with more than one recorded measurement for this model — every one shown, not just the headline number.
6 facts cross-checked across data sources: 1 corroborated, 1 single-source, 4 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: curated_benchmark_scores, curated_capability_claims, epoch_benchmarks, huggingface_models +3 more
DeepSeek-R1 is an AI model developed by DeepSeek, released 2025-01-20. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
DeepSeek-R1 has recorded scores on 17 benchmarks, each shown with its evidence status.
DeepSeek-R1 has recorded scores on 18 benchmarks — Lech Mazur Writing, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, SimpleBench, and 12 more. The full table above shows each score with its evidence status.
3 of 18 recorded scores (17%) are independently reproduced rather than self-reported by the lab.
DeepSeek-R1 has 17 tracked claims: 2 independently reproduced, 4 self-reported, 10 unverified, 1 contradicted.
Listed API pricing: $0.28 per million input tokens, $0.42 per million output tokens (prices disputed across sources). See the pricing block for the full breakdown.