DeepSeek·released 2026-08-131 source
DeepSeek V4 Pro 0813 benchmark scores: 11 benchmarks tracked. 64% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 11 benchmarks·best result 98.61 on OTIS Mock AIME 2024-2025 (reproduced)·7 of 11 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 98.61accuracy (%) | reproduced· optimizedT1 | 2026-08-13Epoch AI |
| ARC-AGI | 90.5accuracy (%) | unverified· optimizedT1 | 2026-08-13https://arcprize.org/leaderboard |
| GPQA diamond | 88.89accuracy (%) | reproduced· optimizedT1 | 2026-08-13Epoch AI |
| WeirdML | 66.22accuracy (%) | unverifiedT1 | 2026-08-13https://htihle.github.io/weirdml.html |
| FrontierMath-Tiers-1-3-v2-Private | 64.56accuracy (%) | reproducedT1 | 2026-08-13Epoch AI |
| ARC-AGI-2 | 61.25accuracy (%) | unverifiedT2 | 2026-08-13 |
| SimpleQA Verified | 55.47accuracy (%) | reproducedT1 | 2026-08-13Epoch AI |
| ProofBench | 50accuracy (%) | unverifiedT1 | 2026-08-13https://www.vals.ai/benchmarks/proof_bench |
| Chess Puzzles | 44.23accuracy (%) | reproducedT1 | 2026-08-13Epoch AI |
| Mystery Game Puzzles | 37.21accuracy (%) | reproducedT1 | 2026-08-13Epoch AI |
| FrontierMath-Tier-4-v2-Private | 26.83accuracy (%) | reproducedT1 | 2026-08-13Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by DeepSeek and released in 2026.
2 cited facts
This model is scored on 11 tracked benchmarks. It currently holds no top score on any of them. Because these scores have been independently reproduced, the record can be trusted as verified.
3 cited facts
On average it trails the top score by 23.86 points, a wide gap that places it well behind the leaders. It holds the top score on none of its benchmarks, so it has no category leadership yet.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
DeepSeek V4 Pro 0813 is an AI model developed by DeepSeek, released 2026-08-13. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
DeepSeek V4 Pro 0813 has recorded scores on 11 benchmarks, each shown with its evidence status.
DeepSeek V4 Pro 0813 has recorded scores on 11 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleQA Verified, ARC-AGI, and 5 more. The full table above shows each score with its evidence status.
7 of 11 recorded scores (64%) are independently reproduced rather than self-reported by the lab.
DeepSeek V4 Pro 0813 has 11 tracked claims: 7 independently reproduced, 4 unverified.