Google DeepMind·released 2025-05-061 source
Gemini 2.5 Pro (May 2025) benchmark scores: 10 benchmarks tracked. 20% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 10 benchmarks·best result 95.9 on MATH level 5 (reproduced)·2 of 10 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MATH level 5 | 95.9accuracy (%) | reproduced· optimizedT1 | 2025-05-06Epoch AI |
| GeoBench | 86accuracy (%) | unverifiedT1 | 2025-05-06GeoBench leaderboard |
| Lech Mazur Writing | 80.9accuracy (%) | unverifiedT1 | 2025-05-06lechmazur/writing Github repository |
| Aider polyglot | 76.9accuracy (%) | unverified· optimizedT1 | 2025-05-06Aider LLM Leaderboards |
| Fiction.LiveBench | 66.7accuracy (%) | unverifiedT1 | 2025-05-06Fiction.live leaderboard |
| GPQA diamond | 55.56accuracy (%) | reproduced· optimizedT1 | 2025-05-06Epoch AI |
| DeepResearch Bench | 31.9accuracy (%) | unverifiedT2 | 2025-05-06 |
| The Agent Company | 30.3accuracy (%) | unverifiedT1 | 2025-05-06TheAgentCompany leaderboard |
| HLE | 13.66accuracy (%) | unverifiedT2 | 2025-05-06 |
| VPCT | 10.75accuracy (%) | unverifiedT1 | 2025-05-06VPCT leaderboard |
Claims drawn from cited facts, not live model generation.
This model originates from Google DeepMind and was released in 2025.
2 cited facts
This model is scored on 10 tracked benchmarks and holds no current top score among them. Because the record is independently reproduced, the scores are backed by outside evaluation and can be treated as verified results.
3 cited facts
With an average SOTA gap of 23.03 points, it trails the leaders by a wide margin, placing it well behind the front of the pack rather than at the frontier. It holds the top score on none of the tracked benchmarks, so there is no current category leadership.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Gemini 2.5 Pro (May 2025) is an AI model developed by Google DeepMind, released 2025-05-06. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Gemini 2.5 Pro (May 2025) has recorded scores on 10 benchmarks, each shown with its evidence status.
Gemini 2.5 Pro (May 2025) has recorded scores on 10 benchmarks — Lech Mazur Writing, GPQA diamond, MATH level 5, Aider polyglot, GeoBench, The Agent Company, and 4 more. The full table above shows each score with its evidence status.
2 of 10 recorded scores (20%) are independently reproduced rather than self-reported by the lab.
Gemini 2.5 Pro (May 2025) has 10 tracked claims: 2 independently reproduced, 8 unverified.