Google DeepMind·released 2026-08-131 source
Gemini 3.7 Flash benchmark scores: 10 benchmarks tracked, leading 1. 70% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 1 of 10 benchmarks·best result 97.22 on OTIS Mock AIME 2024-2025 (reproduced)·7 of 10 independently reproduced
Head-to-headGemini 3.7 Flash vs GPT-5.4 ProLeads on: GPQA diamond
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 97.22accuracy (%) | reproduced· optimizedT1 | 2026-08-13Epoch AI |
| GPQA diamond | 93.1accuracy (%) | reproduced· optimizedT1 | 2026-08-13Epoch AI |
| FrontierMath-Tiers-1-3-v2-Private | 71.58accuracy (%) | reproducedT1 | 2026-08-13Epoch AI |
| SimpleQA Verified | 71.2accuracy (%) | reproducedT1 | 2026-08-13Epoch AI |
| DeepSWE | 65.49accuracy (%) | unverifiedT1 | 2026-08-13https://deepswe.datacurve.ai/ |
| ProofBench | 58accuracy (%) | unverifiedT1 | 2026-08-13https://www.vals.ai/benchmarks/proof_bench |
| Chess Puzzles | 44.23accuracy (%) | reproducedT1 | 2026-08-13Epoch AI |
| FrontierCode | 43.59accuracy (%) | unverifiedT1 | 2026-08-13https://cognition.com/frontiercode |
| FrontierMath-Tier-4-v2-Private | 36.59accuracy (%) | reproducedT1 | 2026-08-13Epoch AI |
| Mystery Game Puzzles | 30.6accuracy (%) | reproducedT1 | 2026-08-13Epoch AI |
Claims drawn from cited facts, not live model generation.
The model is scored on 10 tracked benchmarks. It holds a current top score on 1 of those benchmarks. Because the results are independently reproduced, this record is backed by outside evaluation and can be read as verified rather than merely vendor-claimed.
3 cited facts
This model trails the best score by an average of 17.88 points, a wide gap that places it well behind the frontier. It holds the top score on 1 benchmark, but that does not offset the overall deficit.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Gemini 3.7 Flash is an AI model developed by Google DeepMind, released 2026-08-13. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Gemini 3.7 Flash has recorded scores on 10 benchmarks, each shown with its evidence status.
Gemini 3.7 Flash has recorded scores on 10 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, Chess Puzzles, SimpleQA Verified, FrontierMath-Tiers-1-3-v2-Private, FrontierMath-Tier-4-v2-Private, and 4 more. The full table above shows each score with its evidence status.
7 of 10 recorded scores (70%) are independently reproduced rather than self-reported by the lab.
Gemini 3.7 Flash has 10 tracked claims: 7 independently reproduced, 3 unverified.