Verdict
GPT-5.6 Sol leads 17–1 across 18 shared benchmarks.
Claude Opus 4.7 1 · GPT-5.6 Sol 17 · higher isn't always better — check the caveats below.
Reported scores — protocols may differ. How we compare
Reasoning: GPT-5.6 Sol leads 4–0 across 4 shared reasoning benchmarks (largest gap: SimpleQA Verified, 71.6 vs 50.6 — single evaluator).
Unknown: GPT-5.6 Sol leads 3–0 across 3 shared unknown benchmarks (largest gap: Mystery Game Puzzles, 53.73 vs 20.69 — single evaluator).
Coding: Evenly split 1–1 across 2 shared coding benchmarks (largest gap: MirrorCode, 31.11 vs 20 — single evaluator).
Math: GPT-5.6 Sol leads 2–0 across 2 shared math benchmarks (largest gap: FrontierMath-Tier-4-v2-Private, 82.93 vs 31.71 — single evaluator).
Vendor claims — compare against the measured scores above. A claim is what a developer says about its own model, not an independent measurement.
Agentic / Tool Use
“Users report being able to hand off their hardest coding work-the kind that previously needed close supervision-to Opus 4.7 with confidence”
Coding
“Opus 4.7 is a notable improvement on Opus 4.6 in advanced software engineering, with particular gains on the most difficult tasks”
Reasoning
“Second, Opus 4.7 thinks more at higher effort levels, particularly on later turns in agentic settings”
No vendor claims on file for GPT-5.6 Sol.
GPT-5.6 Sol leads 17–1 across 18 shared benchmarks. Claude Opus 4.7 leads 1 benchmark and GPT-5.6 Sol leads 17. Higher isn't always better — see the integrity caveats on each benchmark.
Claude Opus 4.7 and GPT-5.6 Sol have 18 benchmarks in common in our data — those are the only rows where a direct, apples-to-apples comparison is drawn.
GPT-5.6 Sol leads reasoning 4–0. GPT-5.6 Sol leads unknown 3–0.