DeepSeek·released 2026-04-22✓ 2 sources
DeepSeek-V4-Pro benchmark scores: 14 benchmarks tracked. 50% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 14 benchmarks·best result 96.66 on OTIS Mock AIME 2024-2025 (reproduced)·7 of 14 independently reproduced·$1.32/$3.96 per M tokens
Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 96.66accuracy (%) | reproduced· optimizedT1 | 2026-04-24Epoch AI |
| GPQA diamond | 87.88accuracy (%) | reproduced· optimizedT1 | 2026-04-24Epoch AI |
| SWE-Bench verified | 77.64accuracy (%) | reproduced· optimizedT1 | 2026-04-24Epoch AI |
| SimpleQA Verified | 57accuracy (%) | reproducedT1 | 2026-04-24Epoch AI |
| SimpleBench | 53.44accuracy (%) | unverifiedT1 | 2026-04-24https://lmcouncil.ai/benchmarks |
| WeirdML | 48.9accuracy (%) | unverifiedT1 | 2026-04-24https://htihle.github.io/weirdml.html |
| FrontierMath-Tiers-1-3-v2-Private | 45.26accuracy (%) | reproducedT1 | 2026-04-24Epoch AI |
| Surface Evolver Bench | 40accuracy (%) | unverifiedT1 | 2026-04-24https://yhenon.github.io/surface-evolver-llm-eval/ |
| FrontierCode | 17.6accuracy (%) | unverifiedT1 | 2026-04-24https://cognition.com/frontiercode |
| ProofBench | 16accuracy (%) | unverifiedT1 | 2026-04-24https://www.vals.ai/benchmarks/proof_bench |
| Chess Puzzles | 15.82accuracy (%) | reproducedT1 | 2026-04-24Epoch AI |
| CL-bench Life | 13.5accuracy (%) | unverifiedT2 | 2026-04-24 |
| CritPt | 12.86accuracy (%) | unverifiedT2 | 2026-04-24 |
| FrontierMath-Tier-4-v2-Private | 2.44accuracy (%) | reproducedT1 | 2026-04-24Epoch AI |
Claims drawn from cited facts, not live model generation.
Developed by DeepSeek and released in 2026, this model is a frontier-scale system with 1598.84 billion parameters. That scale implies considerable headroom for demanding tasks, but also correspondingly higher computational cost and resource demands.
3 cited facts
This model is scored on 14 tracked benchmarks. It currently holds no top score in any of those benchmarks. Because the reported scores are independently reproduced, outside evaluation backs them up, so the record can be trusted as verified.
3 cited facts
Relative to the best scores on its benchmarks, this model trails by an average of 34.3 points, a wide gap that places it well behind the leaders. It holds the top score on none of the benchmarks counted by sota leader count, so it exhibits no genuine category leadership.
2 cited facts
7 facts cross-checked across data sources: 1 corroborated, 2 single-source, 4 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, huggingface_models, litellm_prices, modelsdev_models +1 more
DeepSeek-V4-Pro is an AI model developed by DeepSeek, released 2026-04-22. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
DeepSeek-V4-Pro has recorded scores on 14 benchmarks, each shown with its evidence status.
DeepSeek-V4-Pro has recorded scores on 14 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleBench, SimpleQA Verified, and 8 more. The full table above shows each score with its evidence status.
7 of 14 recorded scores (50%) are independently reproduced rather than self-reported by the lab.
DeepSeek-V4-Pro has 14 tracked claims: 7 independently reproduced, 7 unverified.
Listed API pricing: $1.32 per million input tokens, $3.96 per million output tokens (prices disputed across sources). See the pricing block for the full breakdown.