GLM-5.2 vs Qwen 3.6 Max (Preview)
Verdict
GLM-5.2 leads 3–2 across 5 shared benchmarks.
GLM-5.2 3 · Qwen 3.6 Max (Preview) 2 · higher isn't always better — see caveats.
Per-benchmark head-to-head
Reported scores — protocols may differ. How we compare
Sort
GLM-5.2Qwen 3.6 Max (Preview)
Capabilities
Reasoning
Reasoning: Evenly split 1–1 across 2 shared reasoning benchmarks (largest gap: SimpleQA Verified, 56.93 vs 38.1 — single evaluator).
4 reproduced
What the labs claim
Vendor claims — compare against the measured scores above. A claim is what a developer says about its own model, not an independent measurement.
GLM-5.2
Coding
“Advanced Coding with Flexible Effort: Stronger coding capabilities with multiple thinking effort levels to balance performance and latency”
huggingface.co2026
Context Handling
“We propose IndexShare, which reuses the same indexer across every four sparse attention layers, reducing per-token FLOPs by 2.9x at a 1M context length.”
huggingface.co2026
General
“Pure Open: An MIT open-source license - no regional limits, technical access without borders”
huggingface.co2026
Qwen 3.6 Max (Preview)
Agentic / Tool Use
Coding
“It achieves the top score on six major coding benchmarks - SWE-bench Pro, Terminal-Bench 2.0, SkillsBench, QwenClawBench, QwenWebBench, and SciCode - with substantial gains over its predecessor”
qwen.ai2026