Verdict
Gemini 3.1 Pro leads 13–8 across 22 shared benchmarks.
Claude Opus 4.6 8 · Gemini 3.1 Pro 13 · 1 tied · higher isn't always better — check the caveats below.
Reported scores — protocols may differ. How we compare
Reasoning: Gemini 3.1 Pro leads 5–0 across 5 shared reasoning benchmarks (largest gap: SimpleQA Verified, 77.3 vs 46.49 — single evaluator).
Unknown: Evenly split 2–2 across 4 shared unknown benchmarks (largest gap: ProofBench, 50 vs 26 — single evaluator).
Context Learning: Evenly split 1–1 across 2 shared context learning benchmarks (largest gap: CL-bench, 20.8 vs 20.7 — protocols undisclosed).
Math: Gemini 3.1 Pro leads 1–0 across 2 shared math benchmarks (largest gap: OTIS Mock AIME 2024-2025, 95.6 vs 94.44 — single evaluator).
Vendor claims — compare against the measured scores above. A claim is what a developer says about its own model, not an independent measurement.
Coding
“The new Claude Opus 4.6 improves on its predecessor's coding skills”
Agentic / Tool Use
“Gemini 3.1 Pro was evaluated across a range of benchmarks, including reasoning, multimodal capabilities, agentic tool use, multi-lingual performance, and long-context.”
Multimodal
“a suite of highly capable, natively multimodal reasoning models”
Reasoning
“Gemini 3.1 Pro is Google's most advanced model for complex tasks.”
Gemini 3.1 Pro leads 13–8 across 22 shared benchmarks. Claude Opus 4.6 leads 8 benchmarks and Gemini 3.1 Pro leads 13, with 1 tied. Higher isn't always better — see the integrity caveats on each benchmark.
Claude Opus 4.6 and Gemini 3.1 Pro have 22 benchmarks in common in our data — those are the only rows where a direct, apples-to-apples comparison is drawn.
Gemini 3.1 Pro leads reasoning 5–0. Gemini 3.1 Pro leads math 1–0.