Google DeepMind·released 2025-06-051 source
Gemini 2.5 Pro (Jun 2025) benchmark scores: 24 benchmarks tracked. 29% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 24 benchmarks·best result 91.7 on Fiction.LiveBench (unverified)·7 of 24 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| Fiction.LiveBench | 91.7accuracy (%) | unverifiedT1 | 2025-06-05Fiction.live leaderboard |
| OTIS Mock AIME 2024-2025 | 84.71accuracy (%) | reproduced· optimizedT1 | 2025-06-05Epoch AI |
| Lech Mazur Writing | 83.8accuracy (%) | unverifiedT1 | 2025-06-05lechmazur/writing Github repository |
| Aider polyglot | 83.1accuracy (%) | unverified· optimizedT1 | 2025-06-05Aider LLM Leaderboards |
| GPQA diamond | 80.39accuracy (%) | reproduced· optimizedT1 | 2025-06-05Epoch AI |
| SWE-Bench verified | 57.56accuracy (%) | reproduced· optimizedT1 | 2025-06-05Epoch AI |
| SimpleQA Verified | 56accuracy (%) | reproducedT1 | 2025-06-05Epoch AI |
| METR Time Horizons | 55.44accuracy (%) | unverifiedT1 | 2025-06-05METR - Measuring AI Ability to Complete Long Tasks |
| SimpleBench | 54.88accuracy (%) | unverifiedT1 | 2025-06-05SimpleBench Leaderboard |
| WeirdML | 54.03accuracy (%) | unverifiedT1 | 2025-06-05WeirdML Leaderboard |
| DeepResearch Bench | 49.71accuracy (%) | unverifiedT1 | 2025-06-05DeepResearchBench Leaderboard |
| ARC-AGI | 41accuracy (%) | unverified· optimizedT1 | 2025-06-05ARC Prize Leaderboard |
| Terminal Bench | 32.6accuracy (%) | unverified· optimizedT1 | 2025-06-05Terminal-Bench v2 Leaderboard |
| FrontierMath-Tiers-1-3-v2-Private | 24.56accuracy (%) | reproducedT1 | 2025-06-05Epoch AI |
| GDPval | 23.3accuracy (%) | unverifiedT2 | 2025-06-05 |
| VPCT | 19.6accuracy (%) | unverifiedT1 | 2025-06-05VPCT leaderboard |
| HLE | 17.69accuracy (%) | unverifiedT2 | 2025-06-05 |
| Chess Puzzles | 15.82accuracy (%) | reproducedT1 | 2025-06-05Epoch AI |
| APEX-Agents | 6.6accuracy (%) | unverifiedT2 | 2025-06-05 |
| ARC-AGI-2 | 4.86accuracy (%) | unverifiedT2 | 2025-06-05 |
| GSO-Bench | 3.92accuracy (%) | unverifiedT1 | 2025-06-05GSO Leaderboard |
| CritPt | 2accuracy (%) | unverifiedT2 | 2025-06-05 |
| Remote Labor Index | 0.83accuracy (%) | unverifiedT2 | 2025-06-05 |
| FrontierMath-Tier-4-v2-Private | 0accuracy (%) | reproducedT1 | 2025-06-05Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by Google DeepMind and released in 2025. No parameter count is available, so this model's scale cannot be classified, but its origins point to DeepMind's research lineage.
3 cited facts
This model is scored on 24 tracked benchmarks and currently holds no top score on any of them. Because its results have been independently reproduced, these scores can be trusted as verified.
3 cited facts
On average, this model trails the leading score by 34.19 points on its benchmarks, a wide gap that places it well behind the front-of-pack. It holds the top score on none of the benchmark leaderboards, confirming it is not a category leader.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Gemini 2.5 Pro (Jun 2025) is an AI model developed by Google DeepMind, released 2025-06-05. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Gemini 2.5 Pro (Jun 2025) has recorded scores on 24 benchmarks, each shown with its evidence status.
Gemini 2.5 Pro (Jun 2025) has recorded scores on 24 benchmarks — Lech Mazur Writing, GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleBench, and 18 more. The full table above shows each score with its evidence status.
7 of 24 recorded scores (29%) are independently reproduced rather than self-reported by the lab.
Gemini 2.5 Pro (Jun 2025) has 24 tracked claims: 7 independently reproduced, 17 unverified.