Verdict
Kimi K3 leads 15–1 across 16 shared benchmarks.
GLM-5.2 1 · Kimi K3 15 · higher isn't always better — check the caveats below.
Reported scores — protocols may differ. How we compare
Reasoning: Kimi K3 leads 3–0 across 3 shared reasoning benchmarks (largest gap: ARC-AGI, 94.5 vs 77 — single evaluator).
Coding: Kimi K3 leads 2–0 across 2 shared coding benchmarks (largest gap: DeepSWE, 68.51 vs 43.78 — single evaluator).
Math: Kimi K3 leads 2–0 across 2 shared math benchmarks (largest gap: OTIS Mock AIME 2024-2025, 97.22 vs 86.38 — single evaluator).
Unknown: Kimi K3 leads 2–0 across 2 shared unknown benchmarks (largest gap: ProofBench, 87 vs 35 — single evaluator).
Vendor claims — compare against the measured scores above. A claim is what a developer says about its own model, not an independent measurement.
Coding
“Advanced Coding with Flexible Effort: Stronger coding capabilities with multiple thinking effort levels to balance performance and latency”
Context Handling
“We propose IndexShare, which reuses the same indexer across every four sparse attention layers, reducing per-token FLOPs by 2.9x at a 1M context length.”
General
“Pure Open: An MIT open-source license - no regional limits, technical access without borders”
No vendor claims on file for Kimi K3.
Kimi K3 leads 15–1 across 16 shared benchmarks. GLM-5.2 leads 1 benchmark and Kimi K3 leads 15. Higher isn't always better — see the integrity caveats on each benchmark.
GLM-5.2 and Kimi K3 have 16 benchmarks in common in our data — those are the only rows where a direct, apples-to-apples comparison is drawn.
Kimi K3 leads reasoning 3–0. Kimi K3 leads coding 2–0.