Google DeepMind·released 2026-02-191 source
Gemini 3.1 Pro benchmark scores: 27 benchmarks tracked, leading 2. 37% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 2 of 27 benchmarks·best result 98 on ARC-AGI (unverified)·10 of 27 independently reproduced·$2/$12 per M tokens
Head-to-headGemini 3.1 Pro vs GPT-5.4 ProConsensus: LiteLLM · See every model’s pricing →
Leads on: HLE, SimpleQA Verified
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| ARC-AGI | 98accuracy (%) | unverified· optimizedT1 | 2026-02-19ARC Prize Leaderboard |
| OTIS Mock AIME 2024-2025 | 95.6accuracy (%) | reproduced· optimizedT1 | 2026-02-19Epoch AI |
| GPQA diamond | 92.59accuracy (%) | reproduced· optimizedT1 | 2026-02-19Epoch AI |
| Terminal Bench | 80.2accuracy (%) | unverified· optimizedT1 | 2026-02-19https://www.tbench.ai/leaderboard/terminal-bench/2.0 |
| SimpleQA Verified | 77.3accuracy (%) | reproducedT1 | 2026-02-19Epoch AI |
| ARC-AGI-2 | 77.1accuracy (%) | unverifiedT2 | 2026-02-19 |
| METR Time Horizons | 77.03accuracy (%) | self-reportedT1 | 2026-02-19https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ |
| SWE-Bench verified | 75.62accuracy (%) | reproduced· optimizedT1 | 2026-02-19Epoch AI |
| SimpleBench | 75.52accuracy (%) | unverifiedT1 | 2026-02-19SimpleBench Leaderboard |
| WeirdML | 72.1accuracy (%) | unverifiedT1 | 2026-02-19WeirdML Leaderboard |
| FrontierMath-Tiers-1-3-v2-Private | 59.65accuracy (%) | reproducedT1 | 2026-02-19Epoch AI |
| Balrog | 57accuracy (%) | unverifiedT1 | 2026-02-19Balrog Leaderboard |
| Chess Puzzles | 52.65accuracy (%) | reproducedT1 | 2026-02-19Epoch AI |
| HLE | 43.74accuracy (%) | unverifiedT2 | 2026-02-19 |
| APEX-Agents | 33.5accuracy (%) | unverifiedT2 | 2026-02-19 |
| Mystery Game Puzzles | 27.3accuracy (%) | reproducedT1 | 2026-02-19Epoch AI |
| FrontierMath-Tier-4-v2-Private | 26.83accuracy (%) | reproducedT1 | 2026-02-19Epoch AI |
| ExploitBench | 26.1accuracy (%) | unverifiedT2 | 2026-02-19 |
| ProofBench | 26accuracy (%) | unverifiedT1 | 2026-02-19https://www.vals.ai/benchmarks/proof_bench |
| GSO-Bench | 22.55accuracy (%) | unverifiedT1 | 2026-02-19https://gso-bench.github.io/index.html |
| PostTrainBench | 21.59accuracy (%) | unverifiedT2 | 2026-02-19 |
| CL-bench | 20.8accuracy (%) | unverifiedT2 | 2026-02-19 |
| CritPt | 17.71accuracy (%) | unverifiedT2 | 2026-02-19 |
| CL-bench Life | 16.9accuracy (%) | unverifiedT2 | 2026-02-19 |
| EBR-bench | 14.29accuracy (%) | reproducedT1 | 2026-02-19Epoch AI |
| DeepSWE | 11.75accuracy (%) | unverifiedT1 | 2026-02-19https://deepswe.datacurve.ai/ |
| MirrorCode | 8.89accuracy (%) | reproducedT1 | 2026-02-19Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by Google DeepMind and introduced in 2026, marking its origin with a frontier research lab and a contemporary vintage.
2 cited facts
This model is tracked across 27 benchmarks and currently holds the top score on 2 of them. Because its results have been independently reproduced, these scores can be treated as verified rather than merely vendor-claimed.
3 cited facts
On average, this model trails the state-of-the-art leader by 19.12 points under the disclosed harness, a wide gap that places it well behind the front-of-pack. It nevertheless tops the leaderboard on 2 benchmarks, showing sporadic category leadership rather than consistent frontier status.
2 cited facts
5 facts cross-checked across data sources: 5 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: curated_capability_claims, epoch_benchmarks, litellm_prices
Gemini 3.1 Pro is an AI model developed by Google DeepMind, released 2026-02-19. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Gemini 3.1 Pro has recorded scores on 27 benchmarks, each shown with its evidence status.
Gemini 3.1 Pro has recorded scores on 27 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleBench, METR Time Horizons, and 21 more. The full table above shows each score with its evidence status.
10 of 27 recorded scores (37%) are independently reproduced rather than self-reported by the lab.
Gemini 3.1 Pro has 27 tracked claims: 10 independently reproduced, 1 self-reported, 16 unverified.
Listed API pricing: $2 per million input tokens, $12 per million output tokens. See the pricing block for the full breakdown.