Moonshot AI·released 2025-07-111 source
Kimi K2 (Jul 2025) benchmark scores: 7 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 7 benchmarks·best result 85.6 on Lech Mazur Writing (unverified)·0 of 7 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| Lech Mazur Writing | 85.6accuracy (%) | unverifiedT1 | 2025-07-11lechmazur/writing Github repository |
| Fiction.LiveBench | 61.1accuracy (%) | unverifiedT1 | 2025-07-11Fiction.live leaderboard |
| Aider polyglot | 59.1accuracy (%) | unverified· optimizedT1 | 2025-07-11Aider LLM Leaderboards |
| WeirdML | 39.36accuracy (%) | unverifiedT1 | 2025-07-11WeirdML Leaderboard |
| Terminal Bench | 27.8accuracy (%) | unverified· optimizedT1 | 2025-07-11Terminal-Bench v2 Leaderboard |
| SimpleBench | 11.56accuracy (%) | unverifiedT1 | 2025-07-11SimpleBench Leaderboard |
| GSO-Bench | 4.9accuracy (%) | unverifiedT1 | 2025-07-11GSO Leaderboard |
Claims drawn from cited facts, not live model generation.
This model was developed by Moonshot AI and released in 2025.
2 cited facts
For a record of 7 scored benchmarks and no lead, the counts are context rather than conclusion — check the underlying scores before forming a view. The record is unverified, meaning it has neither been confirmed nor disputed by outside evaluation.
3 cited facts
This model trails the leader by 40.45 points on average, a wide gap, and it holds the top score on none of the benchmarks, indicating it is well behind the front-of-pack.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Kimi K2 (Jul 2025) is an AI model developed by Moonshot AI, released 2025-07-11. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Kimi K2 (Jul 2025) has recorded scores on 7 benchmarks, each shown with its evidence status.
Kimi K2 (Jul 2025) has recorded scores on 7 benchmarks — Lech Mazur Writing, WeirdML, SimpleBench, Aider polyglot, GSO-Bench, Fiction.LiveBench, and 1 more. The full table above shows each score with its evidence status.
0 of 7 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Kimi K2 (Jul 2025) has 7 tracked claims: 7 unverified.