Verdict
Gemini 1.5 Pro (May 2024) leads 4–0 across 4 shared benchmarks.
Gemini 1.5 Pro (May 2024) 4 · phi-3-medium 14B 0 · higher isn't always better — check the caveats below.
Reported scores — protocols may differ. How we compare
Reasoning: Gemini 1.5 Pro (May 2024) leads 2–0 across 2 shared reasoning benchmarks (largest gap: GPQA diamond, 27.82 vs 3.45 — single evaluator).
Gemini 1.5 Pro (May 2024) leads 4–0 across 4 shared benchmarks. Gemini 1.5 Pro (May 2024) leads 4 benchmarks and phi-3-medium 14B leads 0. Higher isn't always better — see the integrity caveats on each benchmark.
Gemini 1.5 Pro (May 2024) and phi-3-medium 14B have 4 benchmarks in common in our data — those are the only rows where a direct, apples-to-apples comparison is drawn.
Gemini 1.5 Pro (May 2024) leads reasoning 2–0.