Google DeepMind·released 2024-12-111 source
Gemini 2.0 Pro benchmark scores: 4 benchmarks tracked. 50% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 4 benchmarks·best result 83.46 on MATH level 5 (reproduced)·2 of 4 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MATH level 5 | 83.46accuracy (%) | reproduced· optimizedT1 | 2024-12-11Epoch AI |
| GPQA diamond | 54.21accuracy (%) | reproduced· optimizedT1 | 2024-12-11Epoch AI |
| Fiction.LiveBench | 41.7accuracy (%) | unverifiedT1 | 2024-12-11Fiction.live leaderboard |
| Aider polyglot | 35.6accuracy (%) | unverified· optimizedT1 | 2024-12-11Aider LLM Leaderboards |
Claims drawn from cited facts, not live model generation.
This model was developed by Google DeepMind and released in 2024.
2 cited facts
Its 4 tracked scores are better read as coordinates than as a verdict — enough to place the model on the map, no more. Of those, it holds no current top score. Because its scores have been independently reproduced, the record is backed by outside evaluation and can be trusted.
3 cited facts
The model trails the leader by an average of 40.29 points, a wide gap that places it well behind the leaders, and it holds the top score on none of the benchmarks.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Gemini 2.0 Pro is an AI model developed by Google DeepMind, released 2024-12-11. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Gemini 2.0 Pro has recorded scores on 4 benchmarks, each shown with its evidence status.
Gemini 2.0 Pro has recorded scores on 4 benchmarks — GPQA diamond, MATH level 5, Aider polyglot, Fiction.LiveBench. The full table above shows each score with its evidence status.
2 of 4 recorded scores (50%) are independently reproduced rather than self-reported by the lab.
Gemini 2.0 Pro has 4 tracked claims: 2 independently reproduced, 2 unverified.