Verdict
GPT-5 leads 7–0 across 7 shared benchmarks.
GPT-5 7 · Kimi K2 (Jul 2025) 0 · higher isn't always better — check the caveats below.
Reported scores — protocols may differ. How we compare
Coding: GPT-5 leads 2–0 across 2 shared coding benchmarks (largest gap: Aider polyglot, 88 vs 59.1 — single evaluator).
GPT-5 leads 7–0 across 7 shared benchmarks. GPT-5 leads 7 benchmarks and Kimi K2 (Jul 2025) leads 0. Higher isn't always better — see the integrity caveats on each benchmark.
GPT-5 and Kimi K2 (Jul 2025) have 7 benchmarks in common in our data — those are the only rows where a direct, apples-to-apples comparison is drawn.
GPT-5 leads coding 2–0.