Verdict
Gemini 3.1 Pro leads 12–2 across 14 shared benchmarks.
Gemini 3.1 Pro 12 · Kimi K2.6 2 · higher isn't always better — check the caveats below.
Reported scores — protocols may differ. How we compare
Unknown: Gemini 3.1 Pro leads 3–0 across 3 shared unknown benchmarks (largest gap: Mystery Game Puzzles, 27.3 vs 9.67 — single evaluator).
Agentic: Gemini 3.1 Pro leads 2–0 across 2 shared agentic benchmarks (largest gap: APEX-Agents, 33.5 vs 18.9 — protocols undisclosed).
Math: Evenly split 1–1 across 2 shared math benchmarks (largest gap: FrontierMath-Tier-4-v2-Private, 26.83 vs 25.64 — single evaluator).
Reasoning: Gemini 3.1 Pro leads 2–0 across 2 shared reasoning benchmarks (largest gap: SimpleQA Verified, 77.3 vs 38.7 — single evaluator).
Vendor claims — compare against the measured scores above. A claim is what a developer says about its own model, not an independent measurement.
Agentic / Tool Use
“Gemini 3.1 Pro was evaluated across a range of benchmarks, including reasoning, multimodal capabilities, agentic tool use, multi-lingual performance, and long-context.”
Multimodal
“a suite of highly capable, natively multimodal reasoning models”
Reasoning
“Gemini 3.1 Pro is Google's most advanced model for complex tasks.”
No vendor claims on file for Kimi K2.6.
Gemini 3.1 Pro leads 12–2 across 14 shared benchmarks. Gemini 3.1 Pro leads 12 benchmarks and Kimi K2.6 leads 2. Higher isn't always better — see the integrity caveats on each benchmark.
Gemini 3.1 Pro and Kimi K2.6 have 14 benchmarks in common in our data — those are the only rows where a direct, apples-to-apples comparison is drawn.
Gemini 3.1 Pro leads unknown 3–0. Gemini 3.1 Pro leads agentic 2–0.