Verdict
GPT-5.5 leads 17–6 across 23 shared benchmarks.
Claude Opus 4.7 6 · GPT-5.5 17 · higher isn't always better — check the caveats below.
Reported scores — protocols may differ. How we compare
Reasoning: GPT-5.5 leads 4–0 across 4 shared reasoning benchmarks (largest gap: SimpleQA Verified, 63.1 vs 50.6 — single evaluator).
Agentic: GPT-5.5 leads 2–1 across 3 shared agentic benchmarks (largest gap: ExploitBench, 47.4 vs 26.5 — protocols undisclosed).
Coding: GPT-5.5 leads 2–1 across 3 shared coding benchmarks (largest gap: MirrorCode, 31.11 vs 10 — single evaluator).
Unknown: GPT-5.5 leads 2–1 across 3 shared unknown benchmarks (largest gap: Mystery Game Puzzles, 51.53 vs 20.69 — single evaluator).
Math: GPT-5.5 leads 2–0 across 2 shared math benchmarks (largest gap: FrontierMath-Tier-4-v2-Private, 72.5 vs 31.71 — single evaluator).
Vendor claims — compare against the measured scores above. A claim is what a developer says about its own model, not an independent measurement.
Agentic / Tool Use
“Users report being able to hand off their hardest coding work-the kind that previously needed close supervision-to Opus 4.7 with confidence”
Coding
“Opus 4.7 is a notable improvement on Opus 4.6 in advanced software engineering, with particular gains on the most difficult tasks”
Reasoning
“Second, Opus 4.7 thinks more at higher effort levels, particularly on later turns in agentic settings”
Agentic / Tool Use
“The gains are especially strong in agentic coding, computer use, knowledge work, and early scientific research”
Coding
“It excels at writing and debugging code, researching online, analyzing data, creating documents and spreadsheets, operating software, and moving across tools until a task is finished.”
General
“We're releasing GPT-5.5, our smartest and most intuitive to use model yet, and the next step toward a new way of getting work done on a computer.”
Speed / Latency
“GPT-5.5 matches GPT-5.4 per-token latency in real-world serving, while performing at a much higher level of intelligence.”
GPT-5.5 leads 17–6 across 23 shared benchmarks. Claude Opus 4.7 leads 6 benchmarks and GPT-5.5 leads 17. Higher isn't always better — see the integrity caveats on each benchmark.
Claude Opus 4.7 and GPT-5.5 have 23 benchmarks in common in our data — those are the only rows where a direct, apples-to-apples comparison is drawn.
GPT-5.5 leads reasoning 4–0. GPT-5.5 leads math 2–0.