Google DeepMind·released 2023-05-171 source
PaLM 2-M benchmark scores: 5 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 5 benchmarks·best result 81.7 on TriviaQA (self-reported)·0 of 5 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| TriviaQA | 81.7accuracy (%) | self-reported· optimizedT1 | 2023-05-17PaLM 2 Technical Report |
| HellaSwag | 78.67accuracy (%) | self-reported· optimizedT1 | 2023-05-17PaLM 2 Technical Report |
| Winogrande | 58.4accuracy (%) | self-reported· optimizedT1 | 2023-05-17PaLM 2 Technical Report |
| ARC AI2 | 53.2accuracy (%) | unverified· optimizedT1 | 2023-05-17The Falcon Series of Open Language Models |
| OpenBookQA | 43.2accuracy (%) | self-reported· optimizedT1 | 2023-05-17PaLM 2 Technical Report |
Claims drawn from cited facts, not live model generation.
This model was developed by Google DeepMind and released in 2023.
2 cited facts
On a 5-benchmark base with no top finish, the honest reading is a placed-but-thin record: enough to locate the model, too little to bound it. Because the results are self-reported, they should be regarded as claims rather than independently verified scores.
3 cited facts
The model trails the leader by an average of 24.46 points, a wide gap that places it well behind the frontier. It does not hold the top score on any benchmark, confirming its position well behind the leaders.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
PaLM 2-M is an AI model developed by Google DeepMind, released 2023-05-17. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
PaLM 2-M has recorded scores on 5 benchmarks, each shown with its evidence status.
PaLM 2-M has recorded scores on 5 benchmarks — OpenBookQA, Winogrande, TriviaQA, ARC AI2, HellaSwag. The full table above shows each score with its evidence status.
0 of 5 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
PaLM 2-M has 5 tracked claims: 4 self-reported, 1 unverified.