Google DeepMind·released 2023-05-171 source
PaLM 2-S benchmark scores: 5 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 5 benchmarks·best result 76 on HellaSwag (self-reported)·0 of 5 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| HellaSwag | 76accuracy (%) | self-reported· optimizedT1 | 2023-05-17PaLM 2 Technical Report |
| TriviaQA | 75.2accuracy (%) | self-reported· optimizedT1 | 2023-05-17PaLM 2 Technical Report |
| Winogrande | 55.8accuracy (%) | self-reported· optimizedT1 | 2023-05-17PaLM 2 Technical Report |
| ARC AI2 | 46.13accuracy (%) | unverified· optimizedT1 | 2023-05-17The Falcon Series of Open Language Models |
| OpenBookQA | 41.6accuracy (%) | self-reported· optimizedT1 | 2023-05-17PaLM 2 Technical Report |
Claims drawn from cited facts, not live model generation.
This model, developed by Google DeepMind, was released in 2023.
2 cited facts
This model is tracked across 5 benchmarks, holds no top scores, and because these scores are self-reported, they should be read as claims rather than verified results.
3 cited facts
Behind by 28.55 points on average is a wide gap in this scoring — but it is an average over the model's own benchmark mix under disclosed harnesses, a summary of distances rather than a statement about ability. Confirmation stops at 'on top nowhere' — any stronger conclusion needs margins, and margins are exactly what a leader count throws away.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
PaLM 2-S is an AI model developed by Google DeepMind, released 2023-05-17. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
PaLM 2-S has recorded scores on 5 benchmarks, each shown with its evidence status.
PaLM 2-S has recorded scores on 5 benchmarks — OpenBookQA, Winogrande, TriviaQA, ARC AI2, HellaSwag. The full table above shows each score with its evidence status.
0 of 5 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
PaLM 2-S has 5 tracked claims: 4 self-reported, 1 unverified.