DeepSeek·released 2025-08-211 source
DeepSeek-V3.1 benchmark scores: 5 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 5 benchmarks·best result 71.12 on DTBench (unverified)·0 of 5 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| DTBench | 71.12accuracy (%) | unverifiedT1 | 2025-08-21https://conceptualreasoning.ai/dtbench |
| Fiction.LiveBench | 52.8accuracy (%) | unverifiedT1 | 2025-08-21Fiction.live leaderboard |
| WeirdML | 38.36accuracy (%) | unverifiedT1 | 2025-08-21https://htihle.github.io/weirdml.html |
| LMCA | 28.58accuracy (%) | unverifiedT1 | 2025-08-21https://conceptualreasoning.ai/lmca |
| SimpleBench | 28accuracy (%) | unverifiedT1 | 2025-08-21SimpleBench Leaderboard |
Claims drawn from cited facts, not live model generation.
This model was developed by DeepSeek and was released in 2025.
2 cited facts
This model is scored on 5 tracked benchmarks, and it currently holds no top score on any of them. Because its record is unverified, its reported results have been neither confirmed nor disputed by outside evaluation, so they should be read with that uncertainty in mind.
3 cited facts
On average, this model trails the SOTA leader by 44.26 points under the disclosed harness, a wide gap that reads as well behind the leaders. It holds the top score on 0 benchmarks, meaning it currently has no benchmark on which it leads outright.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
DeepSeek-V3.1 is an AI model developed by DeepSeek, released 2025-08-21. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
DeepSeek-V3.1 has recorded scores on 5 benchmarks, each shown with its evidence status.
DeepSeek-V3.1 has recorded scores on 5 benchmarks — DTBench, WeirdML, LMCA, SimpleBench, Fiction.LiveBench. The full table above shows each score with its evidence status.
0 of 5 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
DeepSeek-V3.1 has 5 tracked claims: 5 unverified.