DeepSeek·released 2025-05-281 source
DeepSeek-R1 (May 2025) benchmark scores: 13 benchmarks tracked. 31% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 13 benchmarks·best result 96.64 on MATH level 5 (reproduced)·4 of 13 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MATH level 5 | 96.64accuracy (%) | reproduced· optimizedT1 | 2025-05-28Epoch AI |
| Lech Mazur Writing | 81.9accuracy (%) | unverifiedT1 | 2025-05-28lechmazur/writing Github repository |
| Fiction.LiveBench | 75accuracy (%) | unverifiedT1 | 2025-05-28Fiction.live leaderboard |
| Aider polyglot | 71.4accuracy (%) | unverified· optimizedT1 | 2025-05-28Aider LLM Leaderboards |
| GPQA diamond | 68.43accuracy (%) | reproduced· optimizedT1 | 2025-05-28Epoch AI |
| OTIS Mock AIME 2024-2025 | 66.36accuracy (%) | reproduced· optimizedT1 | 2025-05-28Epoch AI |
| METR Time Horizons | 53.78accuracy (%) | unverifiedT1 | 2025-05-28METR - Measuring AI Ability to Complete Long Tasks |
| WeirdML | 41.63accuracy (%) | unverifiedT1 | 2025-05-28WeirdML Leaderboard |
| DeepResearch Bench | 35.1accuracy (%) | unverifiedT1 | 2025-05-28DeepResearchBench Leaderboard |
| SimpleBench | 28.96accuracy (%) | unverifiedT1 | 2025-05-28SimpleBench Leaderboard |
| SimpleQA Verified | 27.4accuracy (%) | reproducedT1 | 2025-05-28Epoch AI |
| ARC-AGI | 21.2accuracy (%) | unverified· optimizedT1 | 2025-05-28ARC Prize Leaderboard |
| ARC-AGI-2 | 1.12accuracy (%) | unverifiedT2 | 2025-05-28 |
Claims drawn from cited facts, not live model generation.
This model was developed by DeepSeek and released in 2025, marking its origin in the DeepSeek research lineage.
2 cited facts
This model is scored on 13 tracked benchmarks. It holds no current top score on those benchmarks. Because the scores have been independently reproduced, outside evaluation backs them up, so the record can be treated as verified.
3 cited facts
With an average SOTA gap of 35.86 points, this model sits well behind the leaders on its benchmarks — a wide gap, not a frontier-level margin. It holds the top score on none of its SOTA leaderboards, so there is no category leadership to point to.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
DeepSeek-R1 (May 2025) is an AI model developed by DeepSeek, released 2025-05-28. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
DeepSeek-R1 (May 2025) has recorded scores on 13 benchmarks, each shown with its evidence status.
DeepSeek-R1 (May 2025) has recorded scores on 13 benchmarks — Lech Mazur Writing, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, SimpleBench, and 7 more. The full table above shows each score with its evidence status.
4 of 13 recorded scores (31%) are independently reproduced rather than self-reported by the lab.
DeepSeek-R1 (May 2025) has 13 tracked claims: 4 independently reproduced, 9 unverified.