Google DeepMind·released 2025-03-251 source
Gemini 2.5 Pro (Mar 2025) benchmark scores: 12 benchmarks tracked. 17% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 12 benchmarks·best result 95.56 on MATH level 5 (reproduced)·2 of 12 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MATH level 5 | 95.56accuracy (%) | reproduced· optimizedT1 | 2025-03-25Epoch AI |
| GeoBench | 81accuracy (%) | unverifiedT1 | 2025-03-25GeoBench leaderboard |
| Lech Mazur Writing | 80.5accuracy (%) | unverifiedT1 | 2025-03-25lechmazur/writing Github repository |
| GPQA diamond | 78.45accuracy (%) | reproduced· optimizedT1 | 2025-03-25Epoch AI |
| Aider polyglot | 72.9accuracy (%) | unverified· optimizedT1 | 2025-03-25Aider LLM Leaderboards |
| Fiction.LiveBench | 66.7accuracy (%) | unverifiedT1 | 2025-03-25Fiction.live leaderboard |
| CadEval | 64accuracy (%) | unverifiedT1 | 2025-03-25CadEval Dashboard |
| Balrog | 43.3accuracy (%) | unverifiedT1 | 2025-03-25Balrog Leaderboard |
| SimpleBench | 41.92accuracy (%) | unverifiedT1 | 2025-03-25SimpleBench Leaderboard |
| ARC-AGI | 33accuracy (%) | unverified· optimizedT1 | 2025-03-25ARC Prize Leaderboard |
| VPCT | 22accuracy (%) | unverifiedT1 | 2025-03-25VPCT leaderboard |
| HLE | 14.03accuracy (%) | unverifiedT2 | 2025-03-25 |
Claims drawn from cited facts, not live model generation.
This model was developed by Google DeepMind and released in 2025.
2 cited facts
Solid breadth at 12 benchmarks and no current lead — together the counts describe a model that competes widely without a headline win. Its record is independently reproduced, meaning outside evaluation backs up the scores, so readers can trust the numbers.
3 cited facts
The model trails the state-of-the-art by an average of 24.72 points, placing it well behind the leaders. It holds the top score on none of the benchmarks, indicating it does not achieve category leadership.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Gemini 2.5 Pro (Mar 2025) is an AI model developed by Google DeepMind, released 2025-03-25. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Gemini 2.5 Pro (Mar 2025) has recorded scores on 12 benchmarks, each shown with its evidence status.
Gemini 2.5 Pro (Mar 2025) has recorded scores on 12 benchmarks — Lech Mazur Writing, GPQA diamond, MATH level 5, CadEval, SimpleBench, Aider polyglot, and 6 more. The full table above shows each score with its evidence status.
2 of 12 recorded scores (17%) are independently reproduced rather than self-reported by the lab.
Gemini 2.5 Pro (Mar 2025) has 12 tracked claims: 2 independently reproduced, 10 unverified.