Moonshot AI·released 2026-02-021 source
Kimi K2.5 benchmark scores: 20 benchmarks tracked. 35% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 20 benchmarks·best result 92.19 on OTIS Mock AIME 2024-2025 (reproduced)·7 of 20 independently reproduced·$0.6/$3 per M tokens
Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 92.19accuracy (%) | reproduced· optimizedT1 | 2026-02-02Epoch AI |
| Fiction.LiveBench | 86.1accuracy (%) | unverifiedT1 | 2026-02-02Fiction.live leaderboard |
| GPQA diamond | 83.47accuracy (%) | reproduced· optimizedT1 | 2026-02-02Epoch AI |
| SWE-Bench verified | 73.76accuracy (%) | reproduced· optimizedT1 | 2026-02-02Epoch AI |
| ARC-AGI | 65.33accuracy (%) | unverified· optimizedT2 | 2026-02-02 |
| OSWorld | 63.3accuracy (%) | unverified· optimizedT1 | 2026-02-02OS World Website |
| FrontierMath-2025-02-28-Private | 48.95accuracy (%) | reproduced· optimizedT1 | 2026-02-02Epoch AI |
| WeirdML | 45.6accuracy (%) | unverifiedT1 | 2026-02-02WeirdML Leaderboard |
| Terminal Bench | 43.2accuracy (%) | unverified· optimizedT2 | 2026-02-02 |
| SimpleBench | 36.16accuracy (%) | unverifiedT1 | 2026-02-02SimpleBench Leaderboard |
| SimpleQA Verified | 33.9accuracy (%) | reproducedT1 | 2026-02-02Epoch AI |
| HLE | 20.56accuracy (%) | unverifiedT2 | 2026-02-02 |
| CL-bench | 19.3accuracy (%) | unverifiedT2 | 2026-02-02 |
| APEX-Agents | 14.4accuracy (%) | unverifiedT2 | 2026-02-02 |
| CL-bench Life | 13.2accuracy (%) | unverifiedT2 | 2026-02-02 |
| ARC-AGI-2 | 11.81accuracy (%) | unverifiedT2 | 2026-02-02 |
| PostTrainBench | 10.26accuracy (%) | unverifiedT2 | 2026-02-02 |
| Chess Puzzles | 7.41accuracy (%) | reproducedT1 | 2026-02-02Epoch AI |
| FrontierMath-Tier-4-2025-07-01-Private | 7accuracy (%) | reproducedT1 | 2026-02-02Epoch AI |
| CritPt | 3.14accuracy (%) | unverifiedT2 | 2026-02-02 |
Claims drawn from cited facts, not live model generation.
This model was developed by Moonshot AI. It was released in 2026.
2 cited facts
This model is scored on 20 tracked benchmarks. It holds no current top score among those benchmarks. Because the record is independently reproduced, outside evaluation backs the scores up, so the results can be trusted as verified.
3 cited facts
On average it trails the top score by 28.24 points, a wide gap that places it well behind the leaders. It holds the top score on none of the benchmarks, so it has no category leadership yet.
2 cited facts
6 facts cross-checked across data sources: 1 corroborated, 1 single-source, 4 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: curated_capability_claims, epoch_benchmarks, litellm_prices, modelsdev_models +1 more
Kimi K2.5 is an AI model developed by Moonshot AI, released 2026-02-02. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Kimi K2.5 has recorded scores on 20 benchmarks, each shown with its evidence status.
Kimi K2.5 has recorded scores on 20 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleBench, FrontierMath-2025-02-28-Private, and 14 more. The full table above shows each score with its evidence status.
7 of 20 recorded scores (35%) are independently reproduced rather than self-reported by the lab.
Kimi K2.5 has 20 tracked claims: 7 independently reproduced, 13 unverified.
Listed API pricing: $0.6 per million input tokens, $3 per million output tokens (prices disputed across sources). See the pricing block for the full breakdown.