Google DeepMind·released 2024-06-241 source
Gemma 2 27B benchmark scores: 4 benchmarks tracked. 75% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 4 benchmarks·best result 67.6 on MMLU (unverified)·3 of 4 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MMLU | 67.6accuracy (%) | unverified· optimizedT1 | 2024-06-24Stanford CRFM Leaderboard |
| MATH level 5 | 27.89accuracy (%) | reproduced· optimizedT1 | 2024-06-24Epoch AI |
| GPQA diamond | 15.32accuracy (%) | reproduced· optimizedT1 | 2024-06-24Epoch AI |
| OTIS Mock AIME 2024-2025 | 1.29accuracy (%) | reproduced· optimizedT1 | 2024-06-24Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by Google DeepMind and released in 2024.
2 cited facts
This model is scored on 4 tracked benchmarks and holds no current top score. Because its record is independently reproduced, you can trust the scores as confirmed by outside evaluation.
3 cited facts
At 65.74 points adrift on average, this sits well behind the leaders — a gap that wide also reflects which benchmarks it was scored on, so read it as a portfolio-level distance, not a per-task fact.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Gemma 2 27B is an AI model developed by Google DeepMind, released 2024-06-24. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Gemma 2 27B has recorded scores on 4 benchmarks, each shown with its evidence status.
Gemma 2 27B has recorded scores on 4 benchmarks — GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, MMLU. The full table above shows each score with its evidence status.
3 of 4 recorded scores (75%) are independently reproduced rather than self-reported by the lab.
Gemma 2 27B has 4 tracked claims: 3 independently reproduced, 1 unverified.