Google DeepMind·released 2026-04-021 source
Gemma 4 31B IT benchmark scores: 7 benchmarks tracked. 57% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 7 benchmarks·best result 73.31 on OTIS Mock AIME 2024-2025 (reproduced)·4 of 7 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 73.31accuracy (%) | reproduced· optimizedT1 | 2026-04-02Epoch AI |
| GPQA diamond | 67.68accuracy (%) | reproduced· optimizedT1 | 2026-04-02Epoch AI |
| WeirdML | 52.26accuracy (%) | unverifiedT1 | 2026-04-02https://htihle.github.io/weirdml.html |
| Surface Evolver Bench | 30.62accuracy (%) | unverifiedT1 | 2026-04-02https://yhenon.github.io/surface-evolver-llm-eval/ |
| SimpleQA Verified | 9.56accuracy (%) | reproducedT1 | 2026-04-02Epoch AI |
| CritPt | 1.43accuracy (%) | unverifiedT2 | 2026-04-02 |
| Chess Puzzles | 0.04accuracy (%) | reproducedT1 | 2026-04-02Epoch AI |
Claims drawn from cited facts, not live model generation.
This model is scored on seven tracked benchmarks, but it does not currently top any of them. These results are independently reproduced, so external evaluation supports the reported numbers.
3 cited facts
With an average gap of 45.27 points behind the leader, this model trails the frontier by a wide margin, placing it well behind the leading systems. It also holds the top score on none of the tracked benchmarks, confirming it has not established category leadership.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Gemma 4 31B IT is an AI model developed by Google DeepMind, released 2026-04-02. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Gemma 4 31B IT has recorded scores on 7 benchmarks, each shown with its evidence status.
Gemma 4 31B IT has recorded scores on 7 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleQA Verified, CritPt, and 1 more. The full table above shows each score with its evidence status.
4 of 7 recorded scores (57%) are independently reproduced rather than self-reported by the lab.
Gemma 4 31B IT has 7 tracked claims: 4 independently reproduced, 3 unverified.