Google DeepMind·released 2024-12-111 source
Gemini 2.0 Flash (Feb 2025) benchmark scores: 10 benchmarks tracked. 40% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 10 benchmarks·best result 82.17 on MATH level 5 (reproduced)·4 of 10 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MATH level 5 | 82.17accuracy (%) | reproduced· optimizedT1 | 2024-12-11Epoch AI |
| GeoBench | 77accuracy (%) | unverifiedT1 | 2024-12-11GeoBench leaderboard |
| Fiction.LiveBench | 61.1accuracy (%) | unverifiedT1 | 2024-12-11Fiction.live leaderboard |
| GPQA diamond | 52.19accuracy (%) | reproduced· optimizedT1 | 2024-12-11Epoch AI |
| OTIS Mock AIME 2024-2025 | 31.04accuracy (%) | reproduced· optimizedT1 | 2024-12-11Epoch AI |
| CadEval | 30accuracy (%) | unverifiedT1 | 2024-12-11CadEval Dashboard |
| WeirdML | 25.77accuracy (%) | unverifiedT1 | 2024-12-11WeirdML Leaderboard |
| The Agent Company | 11.4accuracy (%) | unverifiedT1 | 2024-12-11TheAgentCompany experiment results github |
| FrontierMath-2025-02-28-Private | 3.02accuracy (%) | reproduced· optimizedT1 | 2024-12-11Epoch AI |
| ARC-AGI-2 | 1.3accuracy (%) | unverifiedT2 | 2024-12-11 |
Claims drawn from cited facts, not live model generation.
This model was developed by Google DeepMind and released in 2024.
2 cited facts
Breadth of 10 benchmarks with no first place supports one clean conclusion: consistently measured, currently mid-pack. Because its scores are independently reproduced, readers can trust this record.
3 cited facts
The model trails the leader by an average of 45.93 points on its benchmarks, a wide gap that places it well behind the state-of-the-art. It holds the top score on none of the evaluated benchmarks, confirming it is not a category leader.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Gemini 2.0 Flash (Feb 2025) is an AI model developed by Google DeepMind, released 2024-12-11. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Gemini 2.0 Flash (Feb 2025) has recorded scores on 10 benchmarks, each shown with its evidence status.
Gemini 2.0 Flash (Feb 2025) has recorded scores on 10 benchmarks — GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, CadEval, FrontierMath-2025-02-28-Private, and 4 more. The full table above shows each score with its evidence status.
4 of 10 recorded scores (40%) are independently reproduced rather than self-reported by the lab.
Gemini 2.0 Flash (Feb 2025) has 10 tracked claims: 4 independently reproduced, 6 unverified.