Google DeepMind·released 2026-03-031 source
Gemini 3.1 Flash-Lite benchmark scores: 8 benchmarks tracked. 38% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 8 benchmarks·best result 79.98 on OTIS Mock AIME 2024-2025 (reproduced)·3 of 8 independently reproduced·$0.25/$1.5 per M tokens
Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 79.98accuracy (%) | reproduced· optimizedT1 | 2026-03-03Epoch AI |
| GPQA diamond | 75.76accuracy (%) | reproduced· optimizedT1 | 2026-03-03Epoch AI |
| WeirdML | 52.19accuracy (%) | unverifiedT1 | 2026-03-03WeirdML Leaderboard |
| DeepResearch Bench | 36.39accuracy (%) | unverifiedT1 | 2026-03-03https://drb.futuresearch.ai/#drb |
| Chess Puzzles | 21.09accuracy (%) | reproducedT1 | 2026-03-03Epoch AI |
| APEX-Agents | 13accuracy (%) | unverifiedT2 | 2026-03-03 |
| HLE | 4.03accuracy (%) | unverifiedT2 | 2026-03-03 |
| CritPt | 1.14accuracy (%) | unverifiedT2 | 2026-03-03 |
Claims drawn from cited facts, not live model generation.
This model is a 2026 release from Google DeepMind.
2 cited facts
The model is tracked across eight benchmarks, and it currently holds no top score on any of them. Because its reported scores have been independently reproduced by outside evaluation, this benchmark record should be treated as verified rather than vendor-promised.
3 cited facts
On average, this model trails the state-of-the-art leader by 29.99 points under the disclosed harness, a wide gap that places it well behind the front of the pack. It holds the top score on none of the tracked benchmarks (0 of them), so it has no category leadership yet.
2 cited facts
6 facts cross-checked across data sources: 3 corroborated, 2 single-source, 1 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, modelsdev_models, openrouter_models
Gemini 3.1 Flash-Lite is an AI model developed by Google DeepMind, released 2026-03-03. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Gemini 3.1 Flash-Lite has recorded scores on 8 benchmarks, each shown with its evidence status.
Gemini 3.1 Flash-Lite has recorded scores on 8 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, CritPt, DeepResearch Bench, and 2 more. The full table above shows each score with its evidence status.
3 of 8 recorded scores (38%) are independently reproduced rather than self-reported by the lab.
Gemini 3.1 Flash-Lite has 8 tracked claims: 3 independently reproduced, 5 unverified.
Listed API pricing: $0.25 per million input tokens, $1.5 per million output tokens. See the pricing block for the full breakdown.