Verdict
Claude Fable 5.1 leads 11–3 across 14 shared benchmarks.
Claude Fable 5.1 11 · Gemini 3.1 Pro 3 · higher isn't always better — check the caveats below.
Reported scores — protocols may differ. How we compare
Reasoning: Gemini 3.1 Pro leads 2–1 across 3 shared reasoning benchmarks (largest gap: SimpleQA Verified, 73.5 vs 70.8 — protocols undisclosed).
Unknown: Claude Fable 5.1 leads 3–0 across 3 shared unknown benchmarks (largest gap: ProofBench, 100 vs 26 — single evaluator).
Math: Claude Fable 5.1 leads 2–0 across 2 shared math benchmarks (largest gap: FrontierMath-Tier-4-v2-Private, 87.8 vs 26.83 — protocols undisclosed).
Vendor claims — compare against the measured scores above. A claim is what a developer says about its own model, not an independent measurement.
No vendor claims on file for Claude Fable 5.1.
Agentic / Tool Use
“Gemini 3.1 Pro was evaluated across a range of benchmarks, including reasoning, multimodal capabilities, agentic tool use, multi-lingual performance, and long-context.”
Multimodal
“a suite of highly capable, natively multimodal reasoning models”
Reasoning
“Gemini 3.1 Pro is Google's most advanced model for complex tasks.”
Claude Fable 5.1 leads 11–3 across 14 shared benchmarks. Claude Fable 5.1 leads 11 benchmarks and Gemini 3.1 Pro leads 3. Higher isn't always better — see the integrity caveats on each benchmark.
Claude Fable 5.1 and Gemini 3.1 Pro have 14 benchmarks in common in our data — those are the only rows where a direct, apples-to-apples comparison is drawn.
Claude Fable 5.1 leads unknown 3–0. Claude Fable 5.1 leads math 2–0.