Google DeepMind·released 2026-07-211 source
Gemini 3.5 Flash-Lite benchmark scores: 11 benchmarks tracked. 55% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 11 benchmarks·best result 77.78 on GPQA diamond (reproduced)·6 of 11 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| GPQA diamond | 77.78accuracy (%) | reproduced· optimizedT1 | 2026-07-21Epoch AI |
| OTIS Mock AIME 2024-2025 | 71.08accuracy (%) | reproduced· optimizedT1 | 2026-07-21Epoch AI |
| ARC-AGI | 53.5accuracy (%) | unverified· optimizedT1 | 2026-07-21https://arcprize.org/leaderboard |
| WeirdML | 39accuracy (%) | unverifiedT1 | 2026-07-21https://htihle.github.io/weirdml.html |
| FrontierMath-Tiers-1-3-v2-Private | 25.96accuracy (%) | reproducedT1 | 2026-07-21Epoch AI |
| Chess Puzzles | 17.93accuracy (%) | reproducedT1 | 2026-07-21Epoch AI |
| ProofBench | 13accuracy (%) | unverifiedT1 | 2026-07-21https://www.vals.ai/benchmarks/proof_bench |
| Mystery Game Puzzles | 10.77accuracy (%) | reproducedT1 | 2026-07-21Epoch AI |
| ARC-AGI-2 | 10.28accuracy (%) | unverifiedT2 | 2026-07-21 |
| CritPt | 0accuracy (%) | unverifiedT2 | 2026-07-21 |
| FrontierMath-Tier-4-v2-Private | 0accuracy (%) | reproducedT1 | 2026-07-21Epoch AI |
Claims drawn from cited facts, not live model generation.
Developed by Google DeepMind, this model was released in 2026.
2 cited facts
This model is evaluated on 11 tracked benchmarks and currently holds no top score on any of them. Because these results are independently reproduced, the record can be trusted as verified.
3 cited facts
On average, this model trails the leading score by 52.9 points, a widely wide gap that places it well behind the frontier rather than competitive at the top. It holds the top score on none of the surveyed benchmarks, meaning it has no proven category leadership yet.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Gemini 3.5 Flash-Lite is an AI model developed by Google DeepMind, released 2026-07-21. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Gemini 3.5 Flash-Lite has recorded scores on 11 benchmarks, each shown with its evidence status.
Gemini 3.5 Flash-Lite has recorded scores on 11 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, CritPt, ARC-AGI, and 5 more. The full table above shows each score with its evidence status.
6 of 11 recorded scores (55%) are independently reproduced rather than self-reported by the lab.
Gemini 3.5 Flash-Lite has 11 tracked claims: 6 independently reproduced, 5 unverified.