Google DeepMind·released 2026-05-191 source
Gemini 3.5 Flash benchmark scores: 17 benchmarks tracked. 53% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 17 benchmarks·best result 95.55 on OTIS Mock AIME 2024-2025 (reproduced)·9 of 17 independently reproduced·$1.5/$9 per M tokens
Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 95.55accuracy (%) | reproduced· optimizedT1 | 2026-05-19Epoch AI |
| ARC-AGI | 92.5accuracy (%) | unverified· optimizedT1 | 2026-05-19https://arcprize.org/leaderboard |
| GPQA diamond | 90.4accuracy (%) | reproduced· optimizedT1 | 2026-05-19Epoch AI |
| SWE-Bench verified | 79.34accuracy (%) | reproduced· optimizedT1 | 2026-05-19Epoch AI |
| ARC-AGI-2 | 72.08accuracy (%) | unverifiedT2 | 2026-05-19 |
| SimpleBench | 72.04accuracy (%) | unverifiedT1 | 2026-05-19SimpleBench Leaderboard |
| SimpleQA Verified | 68.4accuracy (%) | reproducedT1 | 2026-05-19Epoch AI |
| FrontierMath-Tiers-1-3-v2-Private | 62.81accuracy (%) | reproducedT1 | 2026-05-19Epoch AI |
| WeirdML | 62.64accuracy (%) | unverifiedT1 | 2026-05-19https://htihle.github.io/weirdml.html |
| Surface Evolver Bench | 58.13accuracy (%) | unverifiedT1 | 2026-05-19https://yhenon.github.io/surface-evolver-llm-eval/ |
| Chess Puzzles | 47.39accuracy (%) | reproducedT1 | 2026-05-19Epoch AI |
| DeepSWE | 37.39accuracy (%) | unverifiedT1 | 2026-05-19https://deepswe.datacurve.ai/ |
| ProofBench | 31accuracy (%) | unverifiedT1 | 2026-05-19https://www.vals.ai/benchmarks/proof_bench |
| FrontierMath-Tier-4-v2-Private | 26.83accuracy (%) | reproducedT1 | 2026-05-19Epoch AI |
| Mystery Game Puzzles | 25.09accuracy (%) | reproducedT1 | 2026-05-19Epoch AI |
| CritPt | 13.14accuracy (%) | unverifiedT2 | 2026-05-19 |
| EBR-bench | 4.76accuracy (%) | reproducedT1 | 2026-05-19Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by Google DeepMind and was released in 2026.
2 cited facts
This model is scored on 17 tracked benchmarks and currently holds no top score on any of them. Because the results have been independently reproduced, the record can be treated as externally verified rather than merely vendor-claimed.
3 cited facts
With an average gap of 24.67 points behind the leader, this model sits in a wide-gap band, well behind the front-runners. It holds the top score on none of the benchmarks, so it does not currently show genuine category leadership.
2 cited facts
6 facts cross-checked across data sources: 3 corroborated, 2 single-source, 1 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, modelsdev_models, openrouter_models
Gemini 3.5 Flash is an AI model developed by Google DeepMind, released 2026-05-19. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Gemini 3.5 Flash has recorded scores on 17 benchmarks, each shown with its evidence status.
Gemini 3.5 Flash has recorded scores on 17 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleBench, SimpleQA Verified, and 11 more. The full table above shows each score with its evidence status.
9 of 17 recorded scores (53%) are independently reproduced rather than self-reported by the lab.
Gemini 3.5 Flash has 17 tracked claims: 9 independently reproduced, 8 unverified.
Listed API pricing: $1.5 per million input tokens, $9 per million output tokens. See the pricing block for the full breakdown.