Google DeepMind·released 2025-04-171 source
Gemini 2.5 Flash (Apr 2025) benchmark scores: 10 benchmarks tracked. 10% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 10 benchmarks·best result 76.5 on Lech Mazur Writing (unverified)·1 of 10 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| Lech Mazur Writing | 76.5accuracy (%) | unverifiedT1 | 2025-04-17lechmazur/writing Github repository |
| OTIS Mock AIME 2024-2025 | 73.03accuracy (%) | reproduced· optimizedT1 | 2025-04-17Epoch AI |
| GeoBench | 73accuracy (%) | unverifiedT1 | 2025-04-17GeoBench leaderboard |
| Fiction.LiveBench | 47.2accuracy (%) | unverifiedT1 | 2025-04-17Fiction.live leaderboard |
| Aider polyglot | 47.1accuracy (%) | unverified· optimizedT1 | 2025-04-17Aider LLM Leaderboards |
| WeirdML | 40.95accuracy (%) | unverifiedT1 | 2025-04-17WeirdML Leaderboard |
| ARC-AGI | 32.3accuracy (%) | unverified· optimizedT1 | 2025-04-17ARC Prize Leaderboard |
| DeepResearch Bench | 29.19accuracy (%) | unverifiedT2 | 2025-04-17 |
| HLE | 7.65accuracy (%) | unverifiedT2 | 2025-04-17 |
| VPCT | 7accuracy (%) | unverifiedT1 | 2025-04-17VPCT leaderboard |
Claims drawn from cited facts, not live model generation.
This model was developed by Google DeepMind and released in 2025.
2 cited facts
Depth of 10 benchmarks means the record can absorb an outlier without changing its read — a property thin records lack. It holds no current top score on any of them. Because its scores have been independently reproduced, you can trust the record.
3 cited facts
The model trails the leader by an average of 39.8 points, a wide gap that places it well behind the front-of-pack. It holds the top score on none of the benchmarks, indicating no genuine category leadership.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Gemini 2.5 Flash (Apr 2025) is an AI model developed by Google DeepMind, released 2025-04-17. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Gemini 2.5 Flash (Apr 2025) has recorded scores on 10 benchmarks, each shown with its evidence status.
Gemini 2.5 Flash (Apr 2025) has recorded scores on 10 benchmarks — Lech Mazur Writing, OTIS Mock AIME 2024-2025, WeirdML, Aider polyglot, GeoBench, VPCT, and 4 more. The full table above shows each score with its evidence status.
1 of 10 recorded scores (10%) are independently reproduced rather than self-reported by the lab.
Gemini 2.5 Flash (Apr 2025) has 10 tracked claims: 1 independently reproduced, 9 unverified.