Verdict
Gemini 2.5 Pro (Mar 2025) leads 7–4 across 11 shared benchmarks.
Gemini 2.5 Pro (Mar 2025) 7 · o4-mini 4 · higher isn't always better — check the caveats below.
Reported scores — protocols may differ. How we compare
Reasoning: Gemini 2.5 Pro (Mar 2025) leads 3–1 across 4 shared reasoning benchmarks (largest gap: ARC-AGI, 58.7 vs 33 — single evaluator).
Gemini 2.5 Pro (Mar 2025) leads 7–4 across 11 shared benchmarks. Gemini 2.5 Pro (Mar 2025) leads 7 benchmarks and o4-mini leads 4. Higher isn't always better — see the integrity caveats on each benchmark.
Gemini 2.5 Pro (Mar 2025) and o4-mini have 11 benchmarks in common in our data — those are the only rows where a direct, apples-to-apples comparison is drawn.
Gemini 2.5 Pro (Mar 2025) leads reasoning 3–1.