Moonshot AI·released 2026-07-161 source
Kimi K3 benchmark scores: 17 benchmarks tracked, leading 1. 41% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 1 of 17 benchmarks·best result 97.22 on OTIS Mock AIME 2024-2025 (reproduced)·7 of 17 independently reproduced
Head-to-headKimi K3 vs Claude Fable 5Leads on: Surface Evolver Bench
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 97.22accuracy (%) | reproduced· optimizedT1 | 2026-07-16Epoch AI |
| Surface Evolver Bench | 95accuracy (%) | unverifiedT1 | 2026-07-16https://yhenon.github.io/surface-evolver-llm-eval/ |
| ARC-AGI | 94.5accuracy (%) | unverified· optimizedT1 | 2026-07-16https://arcprize.org/leaderboard |
| GPQA diamond | 90.82accuracy (%) | reproduced· optimizedT1 | 2026-07-16Epoch AI |
| ProofBench | 87accuracy (%) | unverifiedT1 | 2026-07-16https://www.vals.ai/benchmarks/proof_bench |
| WeirdML | 82.57accuracy (%) | unverifiedT1 | 2026-07-16https://htihle.github.io/weirdml.html |
| FrontierMath-Tiers-1-3-v2-Private | 72.18accuracy (%) | reproducedT1 | 2026-07-16Epoch AI |
| DeepSWE | 68.51accuracy (%) | unverifiedT1 | 2026-07-16https://deepswe.datacurve.ai/ |
| ARC-AGI-2 | 60.42accuracy (%) | unverifiedT2 | 2026-07-16 |
| FrontierCode | 44.2accuracy (%) | unverifiedT1 | 2026-07-16https://cognition.com/frontiercode |
| SimpleQA Verified | 42.7accuracy (%) | reproducedT1 | 2026-07-16Epoch AI |
| APEX-Agents | 39.3accuracy (%) | unverifiedT2 | 2026-07-16 |
| FrontierMath-Tier-4-v2-Private | 39.02accuracy (%) | reproducedT1 | 2026-07-16Epoch AI |
| Chess Puzzles | 35.82accuracy (%) | reproducedT1 | 2026-07-16Epoch AI |
| PostTrainBench | 31.96accuracy (%) | unverifiedT2 | 2026-07-16 |
| CritPt | 23.4accuracy (%) | unverifiedT2 | 2026-07-16 |
| Mystery Game Puzzles | 18.48accuracy (%) | reproducedT1 | 2026-07-16Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by Moonshot AI. This model was released in 2026.
2 cited facts
This model is scored on 17 tracked benchmarks and currently holds a top score on 1 of them. Because these results are independently reproduced, the record can be treated as verified rather than merely claimed.
3 cited facts
The model trails the leader by an average of 15.55 points across disclosed benchmarks, a wide gap that places it well behind the front of the pack. It does hold the top score on 1 benchmark, so it is not without a category leadership claim, but the overall average gap signals no sustained frontier position.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Kimi K3 is an AI model developed by Moonshot AI, released 2026-07-16. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Kimi K3 has recorded scores on 17 benchmarks, each shown with its evidence status.
Kimi K3 has recorded scores on 17 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleQA Verified, CritPt, and 11 more. The full table above shows each score with its evidence status.
7 of 17 recorded scores (41%) are independently reproduced rather than self-reported by the lab.
Kimi K3 has 17 tracked claims: 7 independently reproduced, 10 unverified.