Verdict
GPT-4o (May 2024) leads 5–0 across 5 shared benchmarks.
Claude 3 Haiku 0 · GPT-4o (May 2024) 5 · higher isn't always better — check the caveats below.
Reported scores — protocols may differ. How we compare
Math: GPT-4o (May 2024) leads 2–0 across 2 shared math benchmarks (largest gap: MATH level 5, 51.05 vs 14.88 — single evaluator).
GPT-4o (May 2024) leads 5–0 across 5 shared benchmarks. Claude 3 Haiku leads 0 benchmarks and GPT-4o (May 2024) leads 5. Higher isn't always better — see the integrity caveats on each benchmark.
Claude 3 Haiku and GPT-4o (May 2024) have 5 benchmarks in common in our data — those are the only rows where a direct, apples-to-apples comparison is drawn.
GPT-4o (May 2024) leads math 2–0.