Google DeepMind·released 2025-06-171 source
Gemini 2.5 Flash (Jun 2025) benchmark scores: 7 benchmarks tracked. 29% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 7 benchmarks·best result 33.5 on Balrog (unverified)·2 of 7 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| Balrog | 33.5accuracy (%) | unverifiedT1 | 2025-06-17Balrog Leaderboard |
| SimpleBench | 29.44accuracy (%) | unverifiedT1 | 2025-06-17SimpleBench Leaderboard |
| Terminal Bench | 17.1accuracy (%) | unverified· optimizedT2 | 2025-06-17 |
| FrontierMath-2025-02-28-Private | 8.5accuracy (%) | reproduced· optimizedT1 | 2025-06-17Epoch AI |
| FrontierMath-Tier-4-2025-07-01-Private | 6.94accuracy (%) | reproducedT1 | 2025-06-17Epoch AI |
| APEX-Agents | 1.8accuracy (%) | unverifiedT2 | 2025-06-17 |
| CritPt | 1.1accuracy (%) | unverifiedT2 | 2025-06-17 |
Claims drawn from cited facts, not live model generation.
This model was developed by Google DeepMind and released in 2025.
2 cited facts
Evidence here runs 7 benchmarks deep: enough to trust directionally, not enough to settle arguments. It holds no current top score. Because the scores are independently reproduced, outside evaluation backs the record, so readers can trust the reported numbers.
3 cited facts
The model trails the state-of-the-art by an average of 44.01 points, a wide gap that places it well behind the leaders. It does not hold the top score on any benchmark, indicating no category leadership.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Gemini 2.5 Flash (Jun 2025) is an AI model developed by Google DeepMind, released 2025-06-17. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Gemini 2.5 Flash (Jun 2025) has recorded scores on 7 benchmarks, each shown with its evidence status.
Gemini 2.5 Flash (Jun 2025) has recorded scores on 7 benchmarks — SimpleBench, FrontierMath-2025-02-28-Private, Balrog, CritPt, FrontierMath-Tier-4-2025-07-01-Private, APEX-Agents, and 1 more. The full table above shows each score with its evidence status.
2 of 7 recorded scores (29%) are independently reproduced rather than self-reported by the lab.
Gemini 2.5 Flash (Jun 2025) has 7 tracked claims: 2 independently reproduced, 5 unverified.