Google DeepMind·released 2024-05-141 source
Gemini 1.5 Pro (May 2024) benchmark scores: 5 benchmarks tracked, leading 1. 60% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 1 of 5 benchmarks·best result 85.6 on BBH (unverified)·3 of 5 independently reproduced
Head-to-headGemini 1.5 Pro (May 2024) vs DeepSeek-V3Leads on: BBH
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| BBH | 85.6accuracy (%) | unverified· optimizedT1 | 2024-05-14Gemini 1.5 report |
| MMLU | 81.2accuracy (%) | unverified· optimizedT1 | 2024-05-14Gemini 1.5 Report |
| MATH level 5 | 40.75accuracy (%) | reproduced· optimizedT1 | 2024-05-14Epoch AI |
| GPQA diamond | 27.82accuracy (%) | reproduced· optimizedT1 | 2024-05-14Epoch AI |
| OTIS Mock AIME 2024-2025 | 6.71accuracy (%) | reproduced· optimizedT1 | 2024-05-14Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by Google DeepMind and released in 2024.
2 cited facts
Evaluation touches five benchmarks here — a light footprint, so conclusions drawn from it should stay provisional. That one top score is leadership in a single, harness-defined lane: concrete, narrow, and worth checking for comparability before leaning on it. Because the record is independently reproduced, readers can trust these results.
3 cited facts
On average, this model trails the state-of-the-art leader by 43.72 points, a wide gap that places it well behind the front of the pack, though it does hold the top score on one benchmark.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Gemini 1.5 Pro (May 2024) is an AI model developed by Google DeepMind, released 2024-05-14. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Gemini 1.5 Pro (May 2024) has recorded scores on 5 benchmarks, each shown with its evidence status.
Gemini 1.5 Pro (May 2024) has recorded scores on 5 benchmarks — GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, MMLU, BBH. The full table above shows each score with its evidence status.
3 of 5 recorded scores (60%) are independently reproduced rather than self-reported by the lab.
Gemini 1.5 Pro (May 2024) has 5 tracked claims: 3 independently reproduced, 2 unverified.