Verdict
Claude Opus 4.7 leads 12–8 across 20 shared benchmarks.
Claude Opus 4.7 12 · Gemini 3.1 Pro 8 · higher isn't always better — check the caveats below.
Reported scores — protocols may differ. How we compare
Reasoning: Gemini 3.1 Pro leads 5–0 across 5 shared reasoning benchmarks (largest gap: SimpleQA Verified, 77.3 vs 50.6 — single evaluator).
Agentic: Claude Opus 4.7 leads 2–0 across 2 shared agentic benchmarks (largest gap: ExploitBench, 28 vs 26 — protocols undisclosed).
Math: Claude Opus 4.7 leads 2–0 across 2 shared math benchmarks (largest gap: FrontierMath-Tier-4-v2-Private, 31.71 vs 26.83 — single evaluator).
Vendor claims — compare against the measured scores above. A claim is what a developer says about its own model, not an independent measurement.
Agentic / Tool Use
“Users report being able to hand off their hardest coding work-the kind that previously needed close supervision-to Opus 4.7 with confidence”
Coding
“Opus 4.7 is a notable improvement on Opus 4.6 in advanced software engineering, with particular gains on the most difficult tasks”
Reasoning
“Second, Opus 4.7 thinks more at higher effort levels, particularly on later turns in agentic settings”
Agentic / Tool Use
“Gemini 3.1 Pro was evaluated across a range of benchmarks, including reasoning, multimodal capabilities, agentic tool use, multi-lingual performance, and long-context.”
Multimodal
“a suite of highly capable, natively multimodal reasoning models”
Reasoning
“Gemini 3.1 Pro is Google's most advanced model for complex tasks.”
Claude Opus 4.7 leads 12–8 across 20 shared benchmarks. Claude Opus 4.7 leads 12 benchmarks and Gemini 3.1 Pro leads 8. Higher isn't always better — see the integrity caveats on each benchmark.
Claude Opus 4.7 and Gemini 3.1 Pro have 20 benchmarks in common in our data — those are the only rows where a direct, apples-to-apples comparison is drawn.
Gemini 3.1 Pro leads reasoning 5–0. Claude Opus 4.7 leads agentic 2–0.