Google DeepMind·released 2025-12-171 source
Gemini 3 Flash benchmark scores: 19 benchmarks tracked, leading 1. 42% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 1 of 19 benchmarks·best result 95.55 on OTIS Mock AIME 2024-2025 (reproduced)·8 of 19 independently reproduced·$0.5/$3 per M tokens
Head-to-headGemini 3 Flash vs Gemini 2.5 Pro (May 2025)Consensus: LiteLLM · See every model’s pricing →
Leads on: GeoBench
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 95.55accuracy (%) | reproduced· optimizedT1 | 2025-12-17Epoch AI |
| GeoBench | 88accuracy (%) | unverifiedT1 | 2025-12-17GeoBench leaderboard |
| GPQA diamond | 85.86accuracy (%) | reproduced· optimizedT1 | 2025-12-17Epoch AI |
| ARC-AGI | 84.67accuracy (%) | unverified· optimizedT1 | 2025-12-17ARC Prize Leaderboard |
| SWE-Bench verified | 75.41accuracy (%) | reproduced· optimizedT1 | 2025-12-17Epoch AI |
| SimpleQA Verified | 67.4accuracy (%) | reproducedT1 | 2025-12-17Epoch AI |
| Terminal Bench | 64.3accuracy (%) | unverified· optimizedT1 | 2025-12-17Terminal-Bench v2 Leaderboard |
| WeirdML | 61.6accuracy (%) | unverifiedT1 | 2025-12-17WeirdML Leaderboard |
| VPCT | 58.9accuracy (%) | unverifiedT1 | 2025-12-17VPCT leaderboard |
| SimpleBench | 53.32accuracy (%) | unverifiedT1 | 2025-12-17SimpleBench Leaderboard |
| FrontierMath-Tiers-1-3-v2-Private | 51.23accuracy (%) | reproducedT1 | 2025-12-17Epoch AI |
| Balrog | 48.1accuracy (%) | unverifiedT1 | 2025-12-17Balrog Leaderboard |
| Chess Puzzles | 36.87accuracy (%) | reproducedT1 | 2025-12-17Epoch AI |
| ARC-AGI-2 | 33.61accuracy (%) | unverifiedT2 | 2025-12-17 |
| APEX-Agents | 24accuracy (%) | unverifiedT2 | 2025-12-17 |
| Mystery Game Puzzles | 18.48accuracy (%) | reproducedT1 | 2025-12-17Epoch AI |
| FrontierMath-Tier-4-v2-Private | 17.07accuracy (%) | reproducedT1 | 2025-12-17Epoch AI |
| ProofBench | 15accuracy (%) | unverifiedT1 | 2025-12-17https://www.vals.ai/benchmarks/proof_bench |
| GSO-Bench | 9.8accuracy (%) | unverifiedT1 | 2025-12-17GSO Leaderboard |
Claims drawn from cited facts, not live model generation.
This model originates from Google DeepMind and was released in 2025, with no official parameter count disclosed.
2 cited facts
This model is scored on 19 tracked benchmarks and currently holds the top score on 1 of them. Because these results have been independently reproduced, they are confirmed by outside evaluation and can be treated as verified.
3 cited facts
Trailing the leader by an average of 27.8 points, this model sits well behind the front of the pack on its benchmarks. It does hold the top score on 1 benchmark, so it has genuine category leadership there, but that does not offset the overall wide gap.
2 cited facts
5 facts cross-checked across data sources: 5 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices
Gemini 3 Flash is an AI model developed by Google DeepMind, released 2025-12-17. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Gemini 3 Flash has recorded scores on 19 benchmarks, each shown with its evidence status.
Gemini 3 Flash has recorded scores on 19 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleBench, SimpleQA Verified, and 13 more. The full table above shows each score with its evidence status.
8 of 19 recorded scores (42%) are independently reproduced rather than self-reported by the lab.
Gemini 3 Flash has 19 tracked claims: 8 independently reproduced, 11 unverified.
Listed API pricing: $0.5 per million input tokens, $3 per million output tokens. See the pricing block for the full breakdown.