Google DeepMind·released 2025-11-181 source
Gemini 3 Pro benchmark scores: 25 benchmarks tracked, leading 3. 28% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 3 of 25 benchmarks·best result 91.38 on OTIS Mock AIME 2024-2025 (reproduced)·7 of 25 independently reproduced·$2/$12 per M tokens
Head-to-headGemini 3 Pro vs Gemini 3.1 ProConsensus: LiteLLM · See every model’s pricing →
Leads on: Balrog, FrontierMath-Tier-4-2025-07-01-Private, VPCT
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 91.38accuracy (%) | reproduced· optimizedT1 | 2025-11-18Epoch AI |
| GPQA diamond | 90.15accuracy (%) | reproduced· optimizedT1 | 2025-11-18Epoch AI |
| VPCT | 86.5accuracy (%) | unverifiedT1 | 2025-11-18VPCT leaderboard |
| GeoBench | 84accuracy (%) | unverifiedT1 | 2025-11-18GeoBench leaderboard |
| ARC-AGI | 75accuracy (%) | unverified· optimizedT1 | 2025-11-18ARC Prize Leaderboard |
| SWE-Bench verified | 72.93accuracy (%) | reproduced· optimizedT1 | 2025-11-18Epoch AI |
| SimpleQA Verified | 72.9accuracy (%) | reproducedT1 | 2025-11-18Epoch AI |
| SimpleBench | 71.68accuracy (%) | unverifiedT1 | 2025-11-18SimpleBench Leaderboard |
| METR Time Horizons | 70.98accuracy (%) | unverifiedT2 | 2025-11-18 |
| WeirdML | 69.93accuracy (%) | unverifiedT1 | 2025-11-18WeirdML Leaderboard |
| Terminal Bench | 69.4accuracy (%) | unverified· optimizedT1 | 2025-11-18https://www.tbench.ai/leaderboard/terminal-bench/2.0 |
| FrontierMath-2025-02-28-Private | 65.96accuracy (%) | reproduced· optimizedT1 | 2025-11-18Epoch AI |
| Balrog | 58.1accuracy (%) | unverifiedT1 | 2025-11-18Balrog Leaderboard |
| GDPval | 40.3accuracy (%) | unverifiedT2 | 2025-11-18 |
| HLE | 34.37accuracy (%) | unverifiedT2 | 2025-11-18 |
| APEX-Agents | 31.5accuracy (%) | unverifiedT2 | 2025-11-18 |
| FrontierMath-Tier-4-2025-07-01-Private | 31.25accuracy (%) | reproducedT1 | 2025-11-18Epoch AI |
| ARC-AGI-2 | 31.11accuracy (%) | unverifiedT2 | 2025-11-18 |
| Chess Puzzles | 27.4accuracy (%) | reproducedT1 | 2025-11-18Epoch AI |
| ProofBench | 20accuracy (%) | unverifiedT1 | 2025-11-18https://www.vals.ai/benchmarks/proof_bench |
| GSO-Bench | 18.6accuracy (%) | unverifiedT1 | 2025-11-18GSO Leaderboard |
| PostTrainBench | 18.12accuracy (%) | unverifiedT2 | 2025-11-18 |
| CL-bench | 15.8accuracy (%) | unverifiedT2 | 2025-11-18 |
| CritPt | 6.9accuracy (%) | unverifiedT2 | 2025-11-18 |
| Remote Labor Index | 1.25accuracy (%) | unverifiedT2 | 2025-11-18 |
Claims drawn from cited facts, not live model generation.
This model was developed by Google DeepMind and released in 2025.
2 cited facts
This model is scored on 25 tracked benchmarks, and it currently holds the top score on 3 of them. These results are backed by outside evaluation, so the record is independently reproduced and can be trusted as verified.
3 cited facts
Under the disclosed benchmark harness, this model trails the leader by an average of 16.8 points, a wide gap that places it well behind the front of the pack. It does hold the top score on 3 benchmarks, but that doesn't offset the large average gap, so it is not yet a leading model overall.
3 cited facts
6 facts cross-checked across data sources: 6 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: curated_capability_claims, epoch_benchmarks, litellm_prices
Gemini 3 Pro is an AI model developed by Google DeepMind, released 2025-11-18. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Gemini 3 Pro has recorded scores on 25 benchmarks, each shown with its evidence status.
Gemini 3 Pro has recorded scores on 25 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleBench, METR Time Horizons, and 19 more. The full table above shows each score with its evidence status.
7 of 25 recorded scores (28%) are independently reproduced rather than self-reported by the lab.
Gemini 3 Pro has 25 tracked claims: 7 independently reproduced, 18 unverified.
Listed API pricing: $2 per million input tokens, $12 per million output tokens. See the pricing block for the full breakdown.