DeepSeek·released 2024-12-25✓ 2 sources
DeepSeek-V3 benchmark scores: 20 benchmarks tracked, leading 1. 20% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 1 of 19 benchmarks·best result 1134 on Codeforces rating (self-reported)·4 of 20 independently reproduced·$0.27/$1.1 per M tokens
Head-to-headDeepSeek-V3 vs Llama 3.1-405BConsensus: LiteLLM · See every model’s pricing →
Leads on: ARC AI2
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| Codeforces rating | 1134rating | self-reported· optimizedT1 | 2025-01-22DeepSeek AI |
| ARC AI2 | 93.73accuracy (%) | self-reported· optimizedT1 | 2024-12-24DeepSeek-V3 Technical Report |
| MATH-500 | 90.2pass@1 (%) | self-reported· optimizedT1 | 2025-01-22DeepSeek AI |
| HellaSwag | 85.2accuracy (%) | self-reported· optimizedT1 | 2024-12-24DeepSeek-V3 Technical Report |
| BBH | 83.33accuracy (%) | self-reported· optimizedT1 | 2024-12-24DeepSeek-V3 Technical Report |
| MMLU | 82.93accuracy (%) | unverified· optimizedT1 | 2024-12-24Stanford CRFM Leaderboard |
| TriviaQA | 82.9accuracy (%) | self-reported· optimizedT1 | 2024-12-24DeepSeek-V3 Technical Report |
| Winogrande | 70.4accuracy (%) | self-reported· optimizedT1 | 2024-12-24DeepSeek-V3 Technical Report |
| PIQA | 69.4accuracy (%) | self-reported· optimizedT1 | 2024-12-24DeepSeek-V3 Technical Report |
| MATH level 5 | 64.85accuracy (%) | reproduced· optimizedT1 | 2024-12-24Epoch AI |
| GPQA diamond | 59.1accuracy (%) | self-reported· optimizedT1 | 2025-01-22DeepSeek AI |
| Aider polyglot | 48.4accuracy (%) | unverified· optimizedT1 | 2024-12-24Aider LLM Leaderboards |
| METR Time Horizons | 47.36accuracy (%) | unverifiedT1 | 2024-12-24METR - Measuring AI Ability to Complete Long Tasks |
| GPQA diamond | 42.05accuracy (%) | reproduced· optimizedT1 | 2024-12-24Epoch AI |
| AIME | 39.2pass@1 (%) | self-reported· optimizedT1 | 2025-01-22DeepSeek AI |
| LiveCodeBench | 36.2pass@1 (%) | self-reported· optimizedT1 | 2025-01-22DeepSeek AI |
| OTIS Mock AIME 2024-2025 | 15.75accuracy (%) | reproduced· optimizedT1 | 2024-12-24Epoch AI |
| FrontierMath-2025-02-28-Private | 3.02accuracy (%) | reproduced· optimizedT1 | 2024-12-24Epoch AI |
| SimpleBench | 2.68accuracy (%) | unverifiedT1 | 2024-12-24SimpleBench Leaderboard |
| CritPt | 0accuracy (%) | unverifiedT2 | 2024-12-24 |
Claims drawn from cited facts, not live model generation.
Developed by DeepSeek, this model is a frontier-scale system with 684.53 billion parameters, and that scale provides considerable headroom at the cost of high computational expense. It was released in 2024.
3 cited facts
This model is tracked on 19 benchmarks and currently tops 1 of them. However, its record is contradicted: at least one claimed score conflicts with independent measurement, so its self-reported numbers should be treated with real caution.
3 cited facts
On average this model trails the SOTA leader by 73.83 points, a wide gap that places it well behind the front-runners, although it does hold the top score on 1 benchmark.
2 cited facts
Benchmarks with more than one recorded measurement for this model — every one shown, not just the headline number.
6 facts cross-checked across data sources: 1 corroborated, 5 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: curated_benchmark_scores, curated_capability_claims, epoch_benchmarks, huggingface_models +1 more
DeepSeek-V3 is an AI model developed by DeepSeek, released 2024-12-25. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
DeepSeek-V3 has recorded scores on 19 benchmarks, each shown with its evidence status.
DeepSeek-V3 has recorded scores on 20 benchmarks — PIQA, Winogrande, TriviaQA, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, and 14 more. The full table above shows each score with its evidence status.
4 of 20 recorded scores (20%) are independently reproduced rather than self-reported by the lab.
DeepSeek-V3 has 19 tracked claims: 3 independently reproduced, 10 self-reported, 5 unverified, 1 contradicted.
Listed API pricing: $0.27 per million input tokens, $1.1 per million output tokens. See the pricing block for the full breakdown.