Google DeepMind·released 2025-05-201 source
Gemini 2.5 Flash (May 2025) benchmark scores: 8 benchmarks tracked. 13% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 8 benchmarks·best result 77.8 on Fiction.LiveBench (unverified)·1 of 8 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| Fiction.LiveBench | 77.8accuracy (%) | unverifiedT1 | 2025-05-20Fiction.live leaderboard |
| GeoBench | 76accuracy (%) | unverifiedT1 | 2025-05-20GeoBench leaderboard |
| OTIS Mock AIME 2024-2025 | 70.8accuracy (%) | reproduced· optimizedT1 | 2025-05-20Epoch AI |
| Aider polyglot | 55.1accuracy (%) | unverified· optimizedT1 | 2025-05-20Aider LLM Leaderboards |
| WeirdML | 40.95accuracy (%) | unverifiedT1 | 2025-05-20WeirdML Leaderboard |
| ARC-AGI | 33.3accuracy (%) | unverified· optimizedT1 | 2025-05-20ARC Prize Leaderboard |
| HLE | 6.47accuracy (%) | unverifiedT2 | 2025-05-20 |
| ARC-AGI-2 | 2.54accuracy (%) | unverifiedT2 | 2025-05-20 |
Claims drawn from cited facts, not live model generation.
This model was developed by Google DeepMind. It was released in 2025.
2 cited facts
On 8 benchmarks without a lead, the fair summary is measured and mid-field — a placement, not a performance ceiling. Its record is independently reproduced, meaning outside evaluation backs up the scores.
3 cited facts
The model trails the state-of-the-art by an average of 40.6 points, a wide gap that places it well behind the leaders. It does not hold the top score on any benchmark, indicating no genuine category leadership.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Gemini 2.5 Flash (May 2025) is an AI model developed by Google DeepMind, released 2025-05-20. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Gemini 2.5 Flash (May 2025) has recorded scores on 8 benchmarks, each shown with its evidence status.
Gemini 2.5 Flash (May 2025) has recorded scores on 8 benchmarks — OTIS Mock AIME 2024-2025, WeirdML, Aider polyglot, GeoBench, ARC-AGI, Fiction.LiveBench, and 2 more. The full table above shows each score with its evidence status.
1 of 8 recorded scores (13%) are independently reproduced rather than self-reported by the lab.
Gemini 2.5 Flash (May 2025) has 8 tracked claims: 1 independently reproduced, 7 unverified.