DeepSeek·released 2024-05-071 source
DeepSeek-V2 (MoE-236B, May 2024) benchmark scores: 7 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 7 benchmarks·best result 89.6 on ARC AI2 (self-reported)·0 of 7 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| ARC AI2 | 89.6accuracy (%) | self-reported· optimizedT1 | 2024-05-07DeepSeek-V3 Technical Report |
| HellaSwag | 82.8accuracy (%) | self-reported· optimizedT1 | 2024-05-07DeepSeek-V3 Technical Report |
| TriviaQA | 80accuracy (%) | self-reported· optimizedT1 | 2024-05-07DeepSeek-V3 Technical Report |
| Winogrande | 72.6accuracy (%) | self-reported· optimizedT1 | 2024-05-07DeepSeek-V3 Technical Report |
| BBH | 71.73accuracy (%) | self-reported· optimizedT1 | 2024-05-07DeepSeek-V3 Technical Report |
| MMLU | 71.2accuracy (%) | self-reported· optimizedT1 | 2024-05-07DeepSeek-V3 Technical Report |
| PIQA | 67.8accuracy (%) | self-reported· optimizedT1 | 2024-05-07DeepSeek-V3 Technical Report |
Claims drawn from cited facts, not live model generation.
This model was developed by DeepSeek and released in 2024.
2 cited facts
This model is scored on 7 tracked benchmarks and currently holds no top scores; because the scores are self-reported, they should be treated as vendor claims rather than independently verified results.
3 cited facts
The model trails the SOTA leader by an average of 9.51 points, indicating a wide gap, and it does not hold the top score on any benchmark.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
DeepSeek-V2 (MoE-236B, May 2024) is an AI model developed by DeepSeek, released 2024-05-07. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
DeepSeek-V2 (MoE-236B, May 2024) has recorded scores on 7 benchmarks, each shown with its evidence status.
DeepSeek-V2 (MoE-236B, May 2024) has recorded scores on 7 benchmarks — PIQA, Winogrande, TriviaQA, MMLU, ARC AI2, BBH, and 1 more. The full table above shows each score with its evidence status.
0 of 7 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
DeepSeek-V2 (MoE-236B, May 2024) has 7 tracked claims: 7 self-reported.