Google DeepMind·released 2024-09-241 source
Gemini 1.5 Pro (Sept 2024) benchmark scores: 11 benchmarks tracked. 27% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 11 benchmarks·best result 82.53 on MMLU (unverified)·3 of 11 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MMLU | 82.53accuracy (%) | unverified· optimizedT1 | 2024-09-24Stanford CRFM Leaderboard |
| MATH level 5 | 70.39accuracy (%) | reproduced· optimizedT1 | 2024-09-24Epoch AI |
| GPQA diamond | 42.97accuracy (%) | reproduced· optimizedT1 | 2024-09-24Epoch AI |
| CadEval | 34accuracy (%) | unverifiedT1 | 2024-09-24CadEval Dashboard |
| OTIS Mock AIME 2024-2025 | 22.98accuracy (%) | reproduced· optimizedT1 | 2024-09-24Epoch AI |
| WeirdML | 22.2accuracy (%) | unverifiedT1 | 2024-09-24WeirdML Leaderboard |
| Balrog | 21accuracy (%) | unverifiedT1 | 2024-09-24Balrog Leaderboard |
| SimpleBench | 12.52accuracy (%) | unverifiedT1 | 2024-09-24SimpleBench Leaderboard |
| The Agent Company | 3.4accuracy (%) | unverifiedT1 | 2024-09-24TheAgentCompany experiment results github |
| ARC-AGI-2 | 0.8accuracy (%) | unverifiedT2 | 2024-09-24 |
| HLE | 0accuracy (%) | unverifiedT2 | 2024-09-24 |
Claims drawn from cited facts, not live model generation.
This model was developed by Google DeepMind and released in 2024.
2 cited facts
An 11-benchmark record with no current lead is a well-supported mid-field placement — the breadth is exactly what makes the placement credible. Because its scores have been independently reproduced by outside evaluators, the record can be trusted.
3 cited facts
On average, this model trails the leader by 48.38 points, a wide gap placing it well behind the front of the pack, and it does not hold the top score on any benchmark.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Gemini 1.5 Pro (Sept 2024) is an AI model developed by Google DeepMind, released 2024-09-24. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Gemini 1.5 Pro (Sept 2024) has recorded scores on 11 benchmarks, each shown with its evidence status.
Gemini 1.5 Pro (Sept 2024) has recorded scores on 11 benchmarks — GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, CadEval, MMLU, and 5 more. The full table above shows each score with its evidence status.
3 of 11 recorded scores (27%) are independently reproduced rather than self-reported by the lab.
Gemini 1.5 Pro (Sept 2024) has 11 tracked claims: 3 independently reproduced, 8 unverified.