Moonshot AI·released 2025-11-061 source
Kimi K2 Thinking benchmark scores: 12 benchmarks tracked. 50% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 12 benchmarks·best result 83.04 on OTIS Mock AIME 2024-2025 (reproduced)·6 of 12 independently reproduced·$0.6/$2.5 per M tokens
Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 83.04accuracy (%) | reproduced· optimizedT1 | 2025-11-06Epoch AI |
| GPQA diamond | 78.96accuracy (%) | reproduced· optimizedT1 | 2025-11-06Epoch AI |
| METR Time Horizons | 59.19accuracy (%) | unverifiedT1 | 2025-11-06METR - Measuring AI Ability to Complete Long Tasks |
| WeirdML | 42.79accuracy (%) | unverifiedT1 | 2025-11-06WeirdML Leaderboard |
| FrontierMath-2025-02-28-Private | 37.55accuracy (%) | reproduced· optimizedT1 | 2025-11-06Epoch AI |
| Terminal Bench | 35.7accuracy (%) | unverified· optimizedT1 | 2025-11-06Terminal-Bench v2 Leaderboard |
| SimpleQA Verified | 31.6accuracy (%) | reproducedT1 | 2025-11-06Epoch AI |
| CL-bench | 17.6accuracy (%) | unverifiedT2 | 2025-11-06 |
| Chess Puzzles | 15.82accuracy (%) | reproducedT1 | 2025-11-06Epoch AI |
| PostTrainBench | 7.25accuracy (%) | unverifiedT2 | 2025-11-06 |
| APEX-Agents | 4.1accuracy (%) | unverifiedT2 | 2025-11-06 |
| FrontierMath-Tier-4-2025-07-01-Private | 0accuracy (%) | reproducedT1 | 2025-11-06Epoch AI |
Claims drawn from cited facts, not live model generation.
Moonshot AI developed this model, and it was released in 2025.
2 cited facts
This model is scored on 12 tracked benchmarks. It holds no current top score on any of those benchmarks. Because the results are independently reproduced, the scores are backed by outside evaluation and can be read as verified.
3 cited facts
On average it trails the state-of-the-art leader by 32.4 points, a wide gap that places it well behind the front-runners. It holds the top score on none of the benchmarks tracked, meaning no genuine category leadership yet.
2 cited facts
6 facts cross-checked across data sources: 3 corroborated, 1 single-source, 2 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, modelsdev_models, openrouter_models
Kimi K2 Thinking is an AI model developed by Moonshot AI, released 2025-11-06. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Kimi K2 Thinking has recorded scores on 12 benchmarks, each shown with its evidence status.
Kimi K2 Thinking has recorded scores on 12 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, METR Time Horizons, FrontierMath-2025-02-28-Private, and 6 more. The full table above shows each score with its evidence status.
6 of 12 recorded scores (50%) are independently reproduced rather than self-reported by the lab.
Kimi K2 Thinking has 12 tracked claims: 6 independently reproduced, 6 unverified.
Listed API pricing: $0.6 per million input tokens, $2.5 per million output tokens. See the pricing block for the full breakdown.