Verdict
Claude Fable 5 leads 12–2 across 14 shared benchmarks.
Claude Fable 5 12 · GPT-5.5 2 · higher isn't always better — check the caveats below.
Reported scores — protocols may differ. How we compare
Coding: Claude Fable 5 leads 2–0 across 2 shared coding benchmarks (largest gap: FrontierCode, 53.5 vs 43 — single evaluator).
Math: Evenly split 1–1 across 2 shared math benchmarks (largest gap: FrontierMath-Tier-4-v2-Private, 87.8 vs 72.5 — single evaluator).
Reasoning: Claude Fable 5 leads 2–0 across 2 shared reasoning benchmarks (largest gap: SimpleBench, 78.28 vs 62.8 — single evaluator).
Unknown: Claude Fable 5 leads 2–0 across 2 shared unknown benchmarks (largest gap: Surface Evolver Bench, 95 vs 88.12 — single evaluator).
Vendor claims — compare against the measured scores above. A claim is what a developer says about its own model, not an independent measurement.
Agentic / Tool Use
“The longer and more complex the task, the larger Fable 5's lead over our other models”
Coding
“It is state-of-the-art on nearly all tested benchmarks of AI capability, showing exceptional performance in software engineering, knowledge work, vision, scientific research, and many other areas”
General
“Fable 5's capabilities exceed those of any model we've ever made generally available”
Multimodal
“It can extract precise numbers from detailed scientific figures and can perform complex vision-based tasks like rebuilding a web app's source code from screenshots alone”
Agentic / Tool Use
“The gains are especially strong in agentic coding, computer use, knowledge work, and early scientific research”
Coding
“It excels at writing and debugging code, researching online, analyzing data, creating documents and spreadsheets, operating software, and moving across tools until a task is finished.”
General
“We're releasing GPT-5.5, our smartest and most intuitive to use model yet, and the next step toward a new way of getting work done on a computer.”
Speed / Latency
“GPT-5.5 matches GPT-5.4 per-token latency in real-world serving, while performing at a much higher level of intelligence.”
Claude Fable 5 leads 12–2 across 14 shared benchmarks. Claude Fable 5 leads 12 benchmarks and GPT-5.5 leads 2. Higher isn't always better — see the integrity caveats on each benchmark.
Claude Fable 5 and GPT-5.5 have 14 benchmarks in common in our data — those are the only rows where a direct, apples-to-apples comparison is drawn.
Claude Fable 5 leads coding 2–0. Claude Fable 5 leads reasoning 2–0.