DeepSeek·released 2026-07-311 source
DeepSeek V4 Flash 0731 benchmark scores: 13 benchmarks tracked. 54% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 13 benchmarks·best result 94.44 on OTIS Mock AIME 2024-2025 (reproduced)·7 of 13 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 94.44accuracy (%) | reproduced· optimizedT1 | 2026-07-31Epoch AI |
| ARC-AGI | 89accuracy (%) | unverified· optimizedT1 | 2026-07-31https://arcprize.org/leaderboard |
| GPQA diamond | 88.05accuracy (%) | reproduced· optimizedT1 | 2026-07-31Epoch AI |
| WeirdML | 62.96accuracy (%) | unverifiedT1 | 2026-07-31https://htihle.github.io/weirdml.html |
| ARC-AGI-2 | 61.39accuracy (%) | unverifiedT2 | 2026-07-31 |
| FrontierMath-Tiers-1-3-v2-Private | 57.54accuracy (%) | reproducedT1 | 2026-07-31Epoch AI |
| ProofBench | 56accuracy (%) | unverifiedT1 | 2026-07-31https://www.vals.ai/benchmarks/proof_bench |
| SimpleQA Verified | 34.67accuracy (%) | reproducedT1 | 2026-07-31Epoch AI |
| Chess Puzzles | 29.5accuracy (%) | reproducedT1 | 2026-07-31Epoch AI |
| Mystery Game Puzzles | 27.3accuracy (%) | reproducedT1 | 2026-07-31Epoch AI |
| FrontierMath-Tier-4-v2-Private | 24.39accuracy (%) | reproducedT1 | 2026-07-31Epoch AI |
| FrontierCode | 18.8accuracy (%) | unverifiedT1 | 2026-07-31https://cognition.com/frontiercode |
| CritPt | 16.57accuracy (%) | unverifiedT2 | 2026-07-31 |
Claims drawn from cited facts, not live model generation.
This model originates from DeepSeek, which developed it. Its release year is 2026.
2 cited facts
This model is scored on 13 tracked benchmarks, and it holds no current top score among them. Because the record is independently reproduced, outside evaluation backs these scores up, so the benchmark results can be treated as verified.
3 cited facts
Based on the disclosed harness, this model trails the leader by 28.57 points on average, a wide gap that places it well behind the front of the pack. It also holds the top score on none of the considered benchmarks, so there is no category leadership yet.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
DeepSeek V4 Flash 0731 is an AI model developed by DeepSeek, released 2026-07-31. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
DeepSeek V4 Flash 0731 has recorded scores on 13 benchmarks, each shown with its evidence status.
DeepSeek V4 Flash 0731 has recorded scores on 13 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleQA Verified, CritPt, and 7 more. The full table above shows each score with its evidence status.
7 of 13 recorded scores (54%) are independently reproduced rather than self-reported by the lab.
DeepSeek V4 Flash 0731 has 13 tracked claims: 7 independently reproduced, 6 unverified.