Google DeepMind·released 2025-01-211 source
Gemini 2.0 Flash Thinking (Jan 2025) benchmark scores: 7 benchmarks tracked. 29% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 7 benchmarks·best result 73.8 on Lech Mazur Writing (unverified)·2 of 7 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| Lech Mazur Writing | 73.8accuracy (%) | unverifiedT1 | 2025-01-21lechmazur/writing Github repository |
| OTIS Mock AIME 2024-2025 | 57.74accuracy (%) | reproduced· optimizedT1 | 2025-01-21Epoch AI |
| Fiction.LiveBench | 52.8accuracy (%) | unverifiedT1 | 2025-01-21Fiction.live leaderboard |
| GPQA diamond | 42.76accuracy (%) | reproduced· optimizedT1 | 2025-01-21Epoch AI |
| Aider polyglot | 18.2accuracy (%) | unverified· optimizedT1 | 2025-01-21Aider LLM Leaderboards |
| SimpleBench | 16.84accuracy (%) | unverifiedT1 | 2025-01-21SimpleBench Leaderboard |
| HLE | 1.85accuracy (%) | unverifiedT2 | 2025-01-21 |
Claims drawn from cited facts, not live model generation.
This model was developed by Google DeepMind and released in 2025.
2 cited facts
Moderate coverage — 7 benchmarks — puts the record between anecdote and profile: usable, with error bars. Standings, not substance: holding no first place records where rivals sit today, and it will re-rank as scores refresh. The scores have been independently reproduced, so the record can be trusted.
3 cited facts
With an average gap of 46.19 points to the state-of-the-art leader and no top scores on any benchmark, this model is well behind the front of the pack.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Gemini 2.0 Flash Thinking (Jan 2025) is an AI model developed by Google DeepMind, released 2025-01-21. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Gemini 2.0 Flash Thinking (Jan 2025) has recorded scores on 7 benchmarks, each shown with its evidence status.
Gemini 2.0 Flash Thinking (Jan 2025) has recorded scores on 7 benchmarks — Lech Mazur Writing, GPQA diamond, OTIS Mock AIME 2024-2025, SimpleBench, Aider polyglot, Fiction.LiveBench, and 1 more. The full table above shows each score with its evidence status.
2 of 7 recorded scores (29%) are independently reproduced rather than self-reported by the lab.
Gemini 2.0 Flash Thinking (Jan 2025) has 7 tracked claims: 2 independently reproduced, 5 unverified.