DeepSeek·released 2025-12-01✓ 2 sources
DeepSeek-V3.2 benchmark scores: 14 benchmarks tracked. 43% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 14 benchmarks·best result 87.81 on OTIS Mock AIME 2024-2025 (reproduced)·6 of 14 independently reproduced·$0.28/$0.4 per M tokens
Consensus: LiteLLM · Cross-check: OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 87.81accuracy (%) | reproduced· optimizedT1 | 2025-12-01Epoch AI |
| GPQA diamond | 77.9accuracy (%) | reproduced· optimizedT1 | 2025-12-01Epoch AI |
| Aider polyglot | 74.2accuracy (%) | unverified· optimizedT1 | 2025-12-01Aider LLM Leaderboards |
| ARC-AGI | 57accuracy (%) | unverified· optimizedT2 | 2025-12-01 |
| Terminal Bench | 39.6accuracy (%) | unverified· optimizedT2 | 2025-12-01 |
| FrontierMath-2025-02-28-Private | 38.77accuracy (%) | reproduced· optimizedT1 | 2025-12-01Epoch AI |
| SimpleQA Verified | 27.5accuracy (%) | reproducedT1 | 2025-12-01Epoch AI |
| CL-bench | 12.4accuracy (%) | unverifiedT2 | 2025-12-01 |
| Chess Puzzles | 9.51accuracy (%) | reproducedT1 | 2025-12-01Epoch AI |
| ProofBench | 8accuracy (%) | unverifiedT1 | 2025-12-01https://www.vals.ai/benchmarks/proof_bench |
| CL-bench Life | 7.4accuracy (%) | unverifiedT2 | 2025-12-01 |
| APEX-Agents | 7accuracy (%) | unverifiedT2 | 2025-12-01 |
| ARC-AGI-2 | 4.03accuracy (%) | unverifiedT2 | 2025-12-01 |
| FrontierMath-Tier-4-2025-07-01-Private | 3.5accuracy (%) | reproducedT1 | 2025-12-01Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by DeepSeek and released in 2025 as a frontier-scale system with 685.4 billion parameters. That scale affords considerable capability headroom, though it carries proportionally high training and inference costs.
4 cited facts
This model is evaluated on 14 tracked benchmarks. It currently holds no top score on any of them. Because the results have been independently reproduced, outside evaluation backs these scores, so the record is reliable.
3 cited facts
With an average gap of 38.24 points behind the leader, this model sits in a wide-gap band, well behind the front of the pack rather than competitive or frontier. It holds the top score on none of its benchmarks, so there is no category leadership yet.
2 cited facts
6 facts cross-checked across data sources: 4 corroborated, 1 single-source, 1 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, huggingface_models, litellm_prices, openrouter_models
DeepSeek-V3.2 is an AI model developed by DeepSeek, released 2025-12-01. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
DeepSeek-V3.2 has recorded scores on 14 benchmarks, each shown with its evidence status.
DeepSeek-V3.2 has recorded scores on 14 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, Chess Puzzles, FrontierMath-2025-02-28-Private, SimpleQA Verified, Aider polyglot, and 8 more. The full table above shows each score with its evidence status.
6 of 14 recorded scores (43%) are independently reproduced rather than self-reported by the lab.
DeepSeek-V3.2 has 14 tracked claims: 6 independently reproduced, 8 unverified.
Listed API pricing: $0.28 per million input tokens, $0.4 per million output tokens. See the pricing block for the full breakdown.