Google DeepMind·released 2026-09-021 source
Gemini 3.8 Flash benchmark scores: 9 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 9 benchmarks·best result 98.89 on OTIS Mock AIME 2024-2025 (unverified)·0 of 9 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 98.89accuracy (%) | unverifiedT2 | 2026-09-02 |
| GPQA diamond | 93.86accuracy (%) | unverifiedT2 | 2026-09-02 |
| DeepSWE | 73.83accuracy (%) | unverifiedT1 | 2026-09-02https://deepswe.datacurve.ai/ |
| SimpleQA Verified | 69.7accuracy (%) | unverifiedT2 | 2026-09-02 |
| FrontierMath-Tiers-1-3-v2-Private | 68.42accuracy (%) | unverifiedT2 | 2026-09-02 |
| Chess Puzzles | 58.96accuracy (%) | unverifiedT2 | 2026-09-02 |
| ProofBench | 48accuracy (%) | unverifiedT1 | 2026-09-02https://www.vals.ai/benchmarks/proof_bench |
| Mystery Game Puzzles | 41.62accuracy (%) | unverifiedT2 | 2026-09-02 |
| FrontierMath-Tier-4-v2-Private | 21.95accuracy (%) | unverifiedT2 | 2026-09-02 |
Claims drawn from cited facts, not live model generation.
This model was developed by Google DeepMind, placing its origin within one of the field's largest and most established research organizations. It was released in 2026, which makes it a recent vintage rather than an older, long-settled generation.
2 cited facts
This model is scored on nine tracked benchmarks, and it currently holds no top score on any of them. Its record is neither confirmed nor disputed by outside evaluation, so the reported scores should be read as unverified rather than as settled results.
3 cited facts
On average this model trails the state-of-the-art leader by 23.67 points across its benchmarks, and under the disclosed harness a gap that wide reads as well behind the leaders rather than competitive or at the frontier. It holds the top score on none of its benchmarks, so there is no category leadership to report here, only the trailing margin itself.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Gemini 3.8 Flash is an AI model developed by Google DeepMind, released 2026-09-02. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Gemini 3.8 Flash has recorded scores on 9 benchmarks, each shown with its evidence status.
Gemini 3.8 Flash has recorded scores on 9 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, SimpleQA Verified, Chess Puzzles, FrontierMath-Tiers-1-3-v2-Private, FrontierMath-Tier-4-v2-Private, and 3 more. The full table above shows each score with its evidence status.
0 of 9 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Gemini 3.8 Flash has 9 tracked claims: 9 unverified.