NVIDIA·released 2026-05-11sources disagree
Kimi K2.6 benchmark scores: 15 benchmarks tracked. 60% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 15 benchmarks·best result 96.11 on OTIS Mock AIME 2024-2025 (reproduced)·9 of 15 independently reproduced·$0.95/$4 per M tokens
Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 96.11accuracy (%) | reproduced· optimizedT1 | 2026-04-20Epoch AI |
| GPQA diamond | 87.71accuracy (%) | reproduced· optimizedT1 | 2026-04-20Epoch AI |
| SWE-Bench verified | 76.65accuracy (%) | reproduced· optimizedT1 | 2026-04-20Epoch AI |
| FrontierMath-Tiers-1-3-v2-Private | 57.19accuracy (%) | reproducedT1 | 2026-04-20Epoch AI |
| WeirdML | 55.86accuracy (%) | unverifiedT1 | 2026-04-20https://htihle.github.io/weirdml.html |
| SimpleQA Verified | 38.7accuracy (%) | reproducedT1 | 2026-04-20Epoch AI |
| FrontierMath-Tier-4-v2-Private | 25.64accuracy (%) | reproducedT1 | 2026-04-20Epoch AI |
| Chess Puzzles | 22.14accuracy (%) | reproducedT1 | 2026-04-20Epoch AI |
| APEX-Agents | 18.9accuracy (%) | unverifiedT2 | 2026-04-20 |
| ExploitBench | 18.4accuracy (%) | unverifiedT2 | 2026-04-20 |
| ProofBench | 16accuracy (%) | unverifiedT1 | 2026-04-20https://www.vals.ai/benchmarks/proof_bench |
| Mystery Game Puzzles | 9.67accuracy (%) | reproducedT1 | 2026-04-20Epoch AI |
| CritPt | 8accuracy (%) | unverifiedT2 | 2026-04-20 |
| OSWorld 2.0 | 4.6accuracy (%) | unverifiedT1 | 2026-04-20https://osworld-v2.xlang.ai/ |
| EBR-bench | 2.38accuracy (%) | reproducedT1 | 2026-04-20Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by NVIDIA. It was released in 2026.
2 cited facts
This model is scored on 15 tracked benchmarks and currently holds no top score on any of them. Because the results have been independently reproduced by outside evaluation, these benchmark numbers can be treated as verified rather than merely vendor-claimed.
3 cited facts
6 facts cross-checked across data sources: 1 corroborated, 5 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, huggingface_models, litellm_prices, modelsdev_models +1 more
Kimi K2.6 is an AI model developed by NVIDIA, released 2026-05-11. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Kimi K2.6 has recorded scores on 15 benchmarks, each shown with its evidence status.
Kimi K2.6 has recorded scores on 15 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleQA Verified, CritPt, and 9 more. The full table above shows each score with its evidence status.
9 of 15 recorded scores (60%) are independently reproduced rather than self-reported by the lab.
Kimi K2.6 has 15 tracked claims: 9 independently reproduced, 6 unverified.
Listed API pricing: $0.95 per million input tokens, $4 per million output tokens (prices disputed across sources). See the pricing block for the full breakdown.