Verdict
GPT-5 leads 13–5 across 18 shared benchmarks.
GPT-5 13 · Grok 4 5 · higher isn't always better — check the caveats below.
Reported scores — protocols may differ. How we compare
Reasoning: Grok 4 leads 3–2 across 5 shared reasoning benchmarks (largest gap: DeepResearch Bench, 55.13 vs 47.9 — single evaluator).
Coding: GPT-5 leads 2–0 across 2 shared coding benchmarks (largest gap: Terminal Bench, 49.6 vs 27.2 — single evaluator).
Game Playing: Evenly split 1–1 across 2 shared game playing benchmarks (largest gap: Balrog, 43.6 vs 32.8 — single evaluator).
GPT-5 leads 13–5 across 18 shared benchmarks. GPT-5 leads 13 benchmarks and Grok 4 leads 5. Higher isn't always better — see the integrity caveats on each benchmark.
GPT-5 and Grok 4 have 18 benchmarks in common in our data — those are the only rows where a direct, apples-to-apples comparison is drawn.
GPT-5 leads coding 2–0. Grok 4 leads reasoning 3–2.