Google DeepMind·released 2026-07-211 source
Gemini 3.6 Flash benchmark scores: 14 benchmarks tracked. 50% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 14 benchmarks·best result 94.16 on OTIS Mock AIME 2024-2025 (reproduced)·7 of 14 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 94.16accuracy (%) | reproduced· optimizedT1 | 2026-07-21Epoch AI |
| GPQA diamond | 92.17accuracy (%) | reproduced· optimizedT1 | 2026-07-21Epoch AI |
| ARC-AGI | 91.17accuracy (%) | unverified· optimizedT1 | 2026-07-21https://arcprize.org/leaderboard |
| SimpleQA Verified | 68.7accuracy (%) | reproducedT1 | 2026-07-21Epoch AI |
| ARC-AGI-2 | 60.42accuracy (%) | unverifiedT2 | 2026-07-21 |
| FrontierMath-Tiers-1-3-v2-Private | 58.95accuracy (%) | reproducedT1 | 2026-07-21Epoch AI |
| WeirdML | 56.1accuracy (%) | unverifiedT1 | 2026-07-21https://htihle.github.io/weirdml.html |
| DeepSWE | 48.56accuracy (%) | unverifiedT1 | 2026-07-21https://deepswe.datacurve.ai/ |
| Chess Puzzles | 40.03accuracy (%) | reproducedT1 | 2026-07-21Epoch AI |
| ProofBench | 36accuracy (%) | unverifiedT1 | 2026-07-21https://www.vals.ai/benchmarks/proof_bench |
| FrontierCode | 34.37accuracy (%) | unverifiedT1 | 2026-07-21https://cognition.com/frontiercode |
| Mystery Game Puzzles | 22.89accuracy (%) | reproducedT1 | 2026-07-21Epoch AI |
| FrontierMath-Tier-4-v2-Private | 21.95accuracy (%) | reproducedT1 | 2026-07-21Epoch AI |
| CritPt | 10.57accuracy (%) | unverifiedT2 | 2026-07-21 |
Claims drawn from cited facts, not live model generation.
Developed by Google DeepMind, this model was released in 2026.
2 cited facts
This model is scored on fourteen tracked benchmarks. It holds no current top score on any of them. The record is independently reproduced, so these scores can be treated as verified.
3 cited facts
With an average gap of 26.4 points to the benchmark leader, this model trails by a wide margin and reads as well behind the front-of-pack, not at the frontier. It holds the top score on zero benchmarks, so it has no category leadership position yet, consistent with that wide gap.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Gemini 3.6 Flash is an AI model developed by Google DeepMind, released 2026-07-21. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Gemini 3.6 Flash has recorded scores on 14 benchmarks, each shown with its evidence status.
Gemini 3.6 Flash has recorded scores on 14 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleQA Verified, CritPt, and 8 more. The full table above shows each score with its evidence status.
7 of 14 recorded scores (50%) are independently reproduced rather than self-reported by the lab.
Gemini 3.6 Flash has 14 tracked claims: 7 independently reproduced, 7 unverified.