Skip to content

GLM-5.2 vs Kimi K2.7 Code

Verdict

GLM-5.2 leads 5–2 across 8 shared benchmarks.

GLM-5.2 5 · Kimi K2.7 Code 2 · 1 tied · higher isn't always better — see caveats.

Per-benchmark head-to-head

Reported scores — protocols may differ. How we compare

Sort
GLM-5.2Kimi K2.7 Code

Capabilities

Math

Math: Evenly split 1–1 across 2 shared math benchmarks (largest gap: FrontierMath-Tier-4-v2-Private, 29.27 vs 12.2 — single evaluator).

4 reproduced

Reasoning

Reasoning: Evenly split 1–1 across 2 shared reasoning benchmarks (largest gap: GPQA diamond, 89.14 vs 86.03 — single evaluator).

4 reproduced

What the labs claim

Vendor claims — compare against the measured scores above. A claim is what a developer says about its own model, not an independent measurement.

GLM-5.2

Coding

  • Advanced Coding with Flexible Effort: Stronger coding capabilities with multiple thinking effort levels to balance performance and latency

Context Handling

  • We propose IndexShare, which reuses the same indexer across every four sparse attention layers, reducing per-token FLOPs by 2.9x at a 1M context length.

General

  • Pure Open: An MIT open-source license - no regional limits, technical access without borders

Kimi K2.7 Code

No vendor claims on file for Kimi K2.7 Code.

A higher number is not always a better model. Each score is task performance under a disclosed harness. Of the 16 measurements across 8 shared benchmarks, 16 single-evaluator or undisclosed-protocol reports; 75% are independently reproduced. GLM-5.2 and Kimi K2.7 Code pricing is cross-check disputed — weigh cost claims with the same caution as benchmark wins. Follow any benchmark to its integrity grade before reading a win as decisive.

Follow the record