Verdict
GPT-5.5 Pro leads 7–2 across 10 shared benchmarks.
GPT-5.5 2 · GPT-5.5 Pro 7 · 1 tied · higher isn't always better — check the caveats below.
Reported scores — protocols may differ. How we compare
Reasoning: GPT-5.5 Pro leads 3–1 across 4 shared reasoning benchmarks (largest gap: SimpleBench, 72.28 vs 62.8 — single evaluator).
Math: GPT-5.5 Pro leads 1–0 across 2 shared math benchmarks (largest gap: FrontierMath-Tier-4-v2-Private, 78.05 vs 72.5 — single evaluator).
Vendor claims — compare against the measured scores above. A claim is what a developer says about its own model, not an independent measurement.
Agentic / Tool Use
“The gains are especially strong in agentic coding, computer use, knowledge work, and early scientific research”
Coding
“It excels at writing and debugging code, researching online, analyzing data, creating documents and spreadsheets, operating software, and moving across tools until a task is finished.”
General
“We're releasing GPT-5.5, our smartest and most intuitive to use model yet, and the next step toward a new way of getting work done on a computer.”
Speed / Latency
“GPT-5.5 matches GPT-5.4 per-token latency in real-world serving, while performing at a much higher level of intelligence.”
General
“In GPT-5.5 Pro, early testers are seeing a significant step up in both the difficulty and quality of work ChatGPT can take on”
Reasoning
“GPT-5.5 Pro, designed for even harder questions and higher-accuracy work, is available to Pro, Business, and Enterprise users.”
Writing
“testers found GPT-5.5 Pro's responses significantly more comprehensive, well-structured, accurate, relevant, and useful”
GPT-5.5 Pro leads 7–2 across 10 shared benchmarks. GPT-5.5 leads 2 benchmarks and GPT-5.5 Pro leads 7, with 1 tied. Higher isn't always better — see the integrity caveats on each benchmark.
GPT-5.5 and GPT-5.5 Pro have 10 benchmarks in common in our data — those are the only rows where a direct, apples-to-apples comparison is drawn.
GPT-5.5 Pro leads reasoning 3–1. GPT-5.5 Pro leads math 1–0.