DeepSeek·released 2026-04-241 source
DeepSeek-V4-Flash benchmark scores: 4 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 4 benchmarks·best result 77.33 on DTBench (unverified)·0 of 4 independently reproduced·$0.3/$1.2 per M tokens
Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| DTBench | 77.33accuracy (%) | unverifiedT1 | 2026-04-24https://conceptualreasoning.ai/dtbench |
| WeirdML | 45.63accuracy (%) | unverifiedT1 | 2026-04-24https://htihle.github.io/weirdml.html |
| LMCA | 42.21accuracy (%) | unverifiedT1 | 2026-04-24https://conceptualreasoning.ai/lmca |
| SimpleBench | 35.56accuracy (%) | unverifiedT1 | 2026-04-24https://lmcouncil.ai/benchmarks |
Claims drawn from cited facts, not live model generation.
This model was developed by DeepSeek. It was released in 2026, placing it among the more recent generation of models.
2 cited facts
This model is scored on 4 tracked benchmarks, and it holds no current top score on any of them. Its record remains unverified: no outside evaluation has either confirmed or disputed the reported scores, so they should be read as neither corroborated nor contradicted.
3 cited facts
The model trails the SOTA leader by an average of 35.56 points, a wide gap under the disclosed harness that reads as well behind the leaders. It holds the top score on 0 benchmarks, meaning no genuine category leadership yet.
2 cited facts
6 facts cross-checked across data sources: 2 single-source, 4 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, modelsdev_models, openrouter_models
DeepSeek-V4-Flash is an AI model developed by DeepSeek, released 2026-04-24. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
DeepSeek-V4-Flash has recorded scores on 4 benchmarks, each shown with its evidence status.
DeepSeek-V4-Flash has recorded scores on 4 benchmarks — DTBench, WeirdML, LMCA, SimpleBench. The full table above shows each score with its evidence status.
0 of 4 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
DeepSeek-V4-Flash has 4 tracked claims: 4 unverified.
Listed API pricing: $0.3 per million input tokens, $1.2 per million output tokens (prices disputed across sources). See the pricing block for the full breakdown.