Google DeepMind·released 2024-02-211 source
Gemma 2B benchmark scores: 8 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 8 benchmarks·best result 61.87 on HellaSwag (unverified)·0 of 8 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| HellaSwag | 61.87accuracy (%) | unverified· optimizedT1 | 2024-02-21Gemma: Open Models Based on Gemini Research and Technology |
| PIQA | 54.6accuracy (%) | unverified· optimizedT1 | 2024-02-21Gemma: Open Models Based on Gemini Research and Technology |
| TriviaQA | 53.2accuracy (%) | unverified· optimizedT1 | 2024-02-21Gemma: Open Models Based on Gemini Research and Technology |
| Winogrande | 30.8accuracy (%) | unverified· optimizedT1 | 2024-02-21Gemma: Open Models Based on Gemini Research and Technology |
| MMLU | 23.07accuracy (%) | unverified· optimizedT1 | 2024-02-21Gemma: Open Models Based on Gemini Research and Technology |
| ARC AI2 | 22.8accuracy (%) | unverified· optimizedT1 | 2024-02-21Gemma: Open Models Based on Gemini Research and Technology |
| GSM8K | 17.7accuracy (%) | unverified· optimizedT1 | 2024-02-21Gemma: Open Models Based on Gemini Research and Technology |
| BBH | 13.6accuracy (%) | unverified· optimizedT1 | 2024-02-21Gemma: Open Models Based on Gemini Research and Technology |
Claims drawn from cited facts, not live model generation.
This model was developed by Google DeepMind and released in 2024.
2 cited facts
This model is scored on 8 tracked benchmarks, but it holds no current top score. The record is unverified, meaning it has neither been confirmed nor disputed by outside evaluation.
3 cited facts
The model trails the SOTA leader by an average of 52.08 points, a wide gap that places it well behind the front-running systems. It holds the top score on none of the benchmarks, confirming it is not a category leader.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Gemma 2B is an AI model developed by Google DeepMind, released 2024-02-21. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Gemma 2B has recorded scores on 8 benchmarks, each shown with its evidence status.
Gemma 2B has recorded scores on 8 benchmarks — PIQA, Winogrande, TriviaQA, MMLU, ARC AI2, BBH, and 2 more. The full table above shows each score with its evidence status.
0 of 8 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Gemma 2B has 8 tracked claims: 8 unverified.