Google DeepMind·released 2024-05-101 source
Gemini 1.5 Flash (Sep 2024) benchmark scores: 9 benchmarks tracked. 44% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 9 benchmarks·best result 76 on GeoBench (unverified)·4 of 9 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| GeoBench | 76accuracy (%) | unverifiedT1 | 2024-05-10GeoBench leaderboard |
| PIQA | 75accuracy (%) | self-reported· optimizedT1 | 2024-05-10Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| MMLU | 65.2accuracy (%) | unverified· optimizedT1 | 2024-05-10Stanford CRFM Leaderboard |
| MATH level 5 | 61.87accuracy (%) | reproduced· optimizedT1 | 2024-05-10Epoch AI |
| GPQA diamond | 29.76accuracy (%) | reproduced· optimizedT1 | 2024-05-10Epoch AI |
| WeirdML | 24.87accuracy (%) | unverifiedT1 | 2024-05-10WeirdML Leaderboard |
| OTIS Mock AIME 2024-2025 | 16.17accuracy (%) | reproduced· optimizedT1 | 2024-05-10Epoch AI |
| Balrog | 14.6accuracy (%) | unverifiedT1 | 2024-05-10Balrog Leaderboard |
| FrontierMath-2025-02-28-Private | 0accuracy (%) | reproduced· optimizedT1 | 2024-05-10Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by Google DeepMind and released in 2024.
2 cited facts
Its coverage of 9 tracked benchmarks is mid-band: broad enough for the overall shape to mean something, not so settled it can't move. The missing top score is the least informative cell on this page; where its scores sit relative to the leaders tells you more than whether it holds a first. Because these results have been independently reproduced by outside evaluators, the record is trustworthy.
3 cited facts
On average, the model trails the leader by 43.67 points. This wide gap places it well behind the frontier.
1 cited fact
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Gemini 1.5 Flash (Sep 2024) is an AI model developed by Google DeepMind, released 2024-05-10. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Gemini 1.5 Flash (Sep 2024) has recorded scores on 9 benchmarks, each shown with its evidence status.
Gemini 1.5 Flash (Sep 2024) has recorded scores on 9 benchmarks — PIQA, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, MMLU, and 3 more. The full table above shows each score with its evidence status.
4 of 9 recorded scores (44%) are independently reproduced rather than self-reported by the lab.
Gemini 1.5 Flash (Sep 2024) has 9 tracked claims: 4 independently reproduced, 1 self-reported, 4 unverified.