DeepSeek·released 2025-03-241 source
DeepSeek-V3 (Mar 2025) benchmark scores: 10 benchmarks tracked. 30% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 10 benchmarks·best result 77 on Lech Mazur Writing (unverified)·3 of 10 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| Lech Mazur Writing | 77accuracy (%) | unverifiedT1 | 2025-03-24lechmazur/writing Github repository |
| MATH level 5 | 75.55accuracy (%) | reproduced· optimizedT1 | 2025-03-24Epoch AI |
| GPQA diamond | 56.82accuracy (%) | reproduced· optimizedT1 | 2025-03-24Epoch AI |
| Aider polyglot | 55.1accuracy (%) | unverified· optimizedT1 | 2025-03-24Aider LLM Leaderboards |
| Fiction.LiveBench | 50accuracy (%) | unverifiedT1 | 2025-03-24Fiction.live leaderboard |
| METR Time Horizons | 49.58accuracy (%) | unverifiedT1 | 2025-03-24METR - Measuring AI Ability to Complete Long Tasks |
| OTIS Mock AIME 2024-2025 | 37.72accuracy (%) | reproduced· optimizedT1 | 2025-03-24Epoch AI |
| WeirdML | 36.08accuracy (%) | unverifiedT1 | 2025-03-24WeirdML Leaderboard |
| SimpleBench | 12.64accuracy (%) | unverifiedT1 | 2025-03-24SimpleBench Leaderboard |
| CritPt | 0accuracy (%) | unverifiedT2 | 2025-03-24 |
Claims drawn from cited facts, not live model generation.
This model is a 2025 release from developer DeepSeek.
2 cited facts
This model is scored on 10 tracked benchmarks and currently holds no top score on any of them. Because these results have been independently reproduced, outside evaluation backs the scores, so the record can be trusted.
3 cited facts
Across its benchmarks, it trails the leader by an average of 39.33 points, a wide gap that places it well behind the front-of-pack. It holds the top score on none of its benchmarks, meaning it has no category leadership yet.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
DeepSeek-V3 (Mar 2025) is an AI model developed by DeepSeek, released 2025-03-24. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
DeepSeek-V3 (Mar 2025) has recorded scores on 10 benchmarks, each shown with its evidence status.
DeepSeek-V3 (Mar 2025) has recorded scores on 10 benchmarks — Lech Mazur Writing, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, SimpleBench, and 4 more. The full table above shows each score with its evidence status.
3 of 10 recorded scores (30%) are independently reproduced rather than self-reported by the lab.
DeepSeek-V3 (Mar 2025) has 10 tracked claims: 3 independently reproduced, 7 unverified.