DeepSeek·released 2025-09-291 source
DeepSeek-V3.2-Exp benchmark scores: 7 benchmarks tracked, leading 1. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 1 of 7 benchmarks·best result 83.3 on Fiction.LiveBench (unverified)·0 of 7 independently reproduced
Head-to-headDeepSeek-V3.2-Exp vs Gemini 2.5 Flash (Sep 2025)Leads on: The Agent Company
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| Fiction.LiveBench | 83.3accuracy (%) | unverifiedT1 | 2025-09-29Fiction.live leaderboard |
| Aider polyglot | 74.2accuracy (%) | unverified· optimizedT1 | 2025-09-29Aider LLM Leaderboards |
| The Agent Company | 42.9accuracy (%) | unverifiedT1 | 2025-09-29TheAgentCompany leaderboard |
| WeirdML | 39.48accuracy (%) | unverifiedT1 | 2025-09-29WeirdML Leaderboard |
| CL-bench | 13.2accuracy (%) | unverifiedT2 | 2025-09-29 |
| CL-bench Life | 9.5accuracy (%) | unverifiedT2 | 2025-09-29 |
| CritPt | 2.86accuracy (%) | unverifiedT2 | 2025-09-29 |
Claims drawn from cited facts, not live model generation.
This model was developed by DeepSeek and released in 2025.
2 cited facts
This model is scored on seven tracked benchmarks. It holds the current top score on one of them. Its results are unverified, so treat the record as neither independently confirmed nor disputed.
3 cited facts
On average across benchmarks, this model trails the leader by 19.57 points, a wide gap that places it well behind the front of the pack under the disclosed harness. It holds the top score on 1 benchmark, indicating genuine category leadership there, but the wide overall average gap shows that is an exception rather than the norm.
2 cited facts
5 facts cross-checked across data sources: 5 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, openrouter_models
DeepSeek-V3.2-Exp is an AI model developed by DeepSeek, released 2025-09-29. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
DeepSeek-V3.2-Exp has recorded scores on 7 benchmarks, each shown with its evidence status.
DeepSeek-V3.2-Exp has recorded scores on 7 benchmarks — WeirdML, Aider polyglot, CritPt, The Agent Company, Fiction.LiveBench, CL-bench, and 1 more. The full table above shows each score with its evidence status.
0 of 7 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
DeepSeek-V3.2-Exp has 7 tracked claims: 7 unverified.