Google DeepMind·released 2024-06-241 source
Gemma 2 9B benchmark scores: 6 benchmarks tracked. 50% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 6 benchmarks·best result 84.9 on GSM8K (self-reported)·3 of 6 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| GSM8K | 84.9accuracy (%) | self-reported· optimizedT1 | 2024-06-24Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| PIQA | 67.4accuracy (%) | self-reported· optimizedT1 | 2024-06-24Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| MMLU | 62.8accuracy (%) | unverified· optimizedT1 | 2024-06-24Stanford CRFM Leaderboard |
| MATH level 5 | 21.01accuracy (%) | reproduced· optimizedT1 | 2024-06-24Epoch AI |
| GPQA diamond | 3.28accuracy (%) | reproduced· optimizedT1 | 2024-06-24Epoch AI |
| OTIS Mock AIME 2024-2025 | 0.46accuracy (%) | reproduced· optimizedT1 | 2024-06-24Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by Google DeepMind and released in 2024.
2 cited facts
Modest on both counts — 6 benchmarks, no leads — this record supports placement, not prediction. The scores are independently reproduced, so the record is backed by outside evaluation.
3 cited facts
This model trails the state-of-the-art by an average of 51.05 points across its benchmarks, placing it well behind the leaders, and it holds the top score on none of them, indicating no genuine category leadership.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Gemma 2 9B is an AI model developed by Google DeepMind, released 2024-06-24. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Gemma 2 9B has recorded scores on 6 benchmarks, each shown with its evidence status.
Gemma 2 9B has recorded scores on 6 benchmarks — PIQA, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, MMLU, GSM8K. The full table above shows each score with its evidence status.
3 of 6 recorded scores (50%) are independently reproduced rather than self-reported by the lab.
Gemma 2 9B has 6 tracked claims: 3 independently reproduced, 2 self-reported, 1 unverified.