Moonshot AI·released 2026-06-121 source
Kimi K2.7 Code benchmark scores: 13 benchmarks tracked. 46% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 13 benchmarks·best result 95.55 on OTIS Mock AIME 2024-2025 (reproduced)·6 of 13 independently reproduced·$0.95/$4 per M tokens
Source: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 95.55accuracy (%) | reproduced· optimizedT1 | 2026-06-12Epoch AI |
| GPQA diamond | 83.84accuracy (%) | reproduced· optimizedT1 | 2026-06-12Epoch AI |
| WeirdML | 54.12accuracy (%) | unverifiedT1 | 2026-06-12https://htihle.github.io/weirdml.html |
| FrontierMath-Tiers-1-3-v2-Private | 54.04accuracy (%) | reproducedT1 | 2026-06-12Epoch AI |
| SimpleBench | 49.48accuracy (%) | unverifiedT1 | 2026-06-12SimpleBench Leaderboard |
| Surface Evolver Bench | 48.75accuracy (%) | unverifiedT1 | 2026-06-12https://yhenon.github.io/surface-evolver-llm-eval/ |
| SimpleQA Verified | 39.2accuracy (%) | reproducedT1 | 2026-06-12Epoch AI |
| DeepSWE | 30.53accuracy (%) | unverifiedT1 | 2026-06-12https://deepswe.datacurve.ai/ |
| FrontierCode | 30.1accuracy (%) | unverifiedT1 | 2026-06-12https://cognition.com/frontiercode |
| APEX-Agents | 27.6accuracy (%) | unverifiedT2 | 2026-06-12 |
| Chess Puzzles | 16.88accuracy (%) | reproducedT1 | 2026-06-12Epoch AI |
| FrontierMath-Tier-4-v2-Private | 12.2accuracy (%) | reproducedT1 | 2026-06-12Epoch AI |
| CritPt | 10accuracy (%) | unverifiedT2 | 2026-06-12 |
Claims drawn from cited facts, not live model generation.
This model was developed by Moonshot AI. It was released in 2026.
2 cited facts
This model is scored on 13 tracked benchmarks and holds no current top score among them. Because the record is independently reproduced, the reported scores are backed by outside evaluation and can be read as verified results.
3 cited facts
Based on an average gap of 32.83 points under the disclosed harness, this model sits well behind the leading scores on its benchmarks. It holds the top score on 0 benchmarks, so it has no category leadership yet.
2 cited facts
5 facts cross-checked across data sources: 1 corroborated, 1 single-source, 3 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, modelsdev_models, openrouter_models
Kimi K2.7 Code is an AI model developed by Moonshot AI, released 2026-06-12. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Kimi K2.7 Code has recorded scores on 13 benchmarks, each shown with its evidence status.
Kimi K2.7 Code has recorded scores on 13 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleBench, SimpleQA Verified, and 7 more. The full table above shows each score with its evidence status.
6 of 13 recorded scores (46%) are independently reproduced rather than self-reported by the lab.
Kimi K2.7 Code has 13 tracked claims: 6 independently reproduced, 7 unverified.
Listed API pricing: $0.95 per million input tokens, $4 per million output tokens (prices disputed across sources). See the pricing block for the full breakdown.