HellaSwag (Harder Endings, Longer contexts, and Low-shot Activities for Situations With Adversarial Generations)
Integrity rank #52 of 61 · 52 models scored · top score 93.73 · GPT-4 (Mar 2023)
Grounded commonsense next-event/sentence completion
Strongest on discrimination, weakest on contamination resistance. Caveat: scores come from an inconsistent mix of harnesses, so head-to-head comparisons are not apples-to-apples.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | GPT-4 (Mar 2023) | 93.73 | unverified· optimizedT1 | 2023-03-15 |
| 2 | Llama 3.1-405B | 85.6 | self-reported· optimizedT1 | 2024-07-23 |
| 3 | Falcon-180B | 85.33 | unverified· optimizedT1 | 2023-09-06 |
| 4 | DeepSeek-V3 | 85.2 | self-reported· optimizedT1 | 2024-12-24 |
| 5 | DeepSeek-V2 (MoE-236B, May 2024) | 82.8 | self-reported· optimizedT1 | 2024-05-07 |
| 6 | PaLM 2-L | 82.4 | self-reported· optimizedT1 | 2023-05-17 |
| 7 | Mixtral 8x7B | 82.27 | unverified· optimizedT1 | 2023-12-11 |
| 8 | Llama 2-70B | 80.4 | self-reported· optimizedT1 | 2023-07-18 |
| 9 | Falcon-40B | 80.37 | self-reported· optimizedT1 | 2023-03-15 |
| 10 | Qwen2.5-72B | 79.73 | self-reported· optimizedT1 | 2024-09-19 |
| 11 | LLaMA-65B | 78.93 | unverified· optimizedT1 | 2023-02-24 |
| 12 | Stable Beluga 2 | 78.8 | self-reported· optimizedT1 | 2023-07-20 |
| 13 | PaLM 2-M | 78.67 | self-reported· optimizedT1 | 2023-05-17 |
| 14 | Qwen2.5-Coder-32B | 77.33 | self-reported· optimizedT1 | 2024-09-18 |
| 15 | Falcon 2 11B | 77.21 | self-reported· optimizedT1 | 2024-05-09 |
| 16 | LLaMA-33B | 77.07 | unverified· optimizedT1 | 2023-02-27 |
| 17 | Nemotron-4 15B | 76.53 | self-reported· optimizedT1 | 2024-02-27 |
| 18 | phi-3-medium 14B | 76.53 | self-reported· optimizedT1 | 2024-04-23 |
| 19 | Gemma 7B | 76.27 | unverified· optimizedT1 | 2024-02-21 |
| 20 | PaLM 2-S | 76 | self-reported· optimizedT1 | 2023-05-17 |
| 21 | Mistral 7B v0.1 | 74.67 | unverified· optimizedT1 | 2023-10-10 |
| 22 | Llama 2-13B | 74.27 | unverified· optimizedT1 | 2023-07-18 |
| 23 | Qwen2.5-Coder-14B | 73.6 | self-reported· optimizedT1 | 2024-09-18 |
| 24 | LLaMA-13B | 72.27 | unverified· optimizedT1 | 2023-02-27 |
| 25 | Falcon-7B | 70.8 | self-reported· optimizedT1 | 2023-04-24 |
| 26 | internlm-20b | 70.8 | self-reported· optimizedT1 | 2023-09-18 |
| 27 | Llama 2-7B | 69.6 | unverified· optimizedT1 | 2023-07-18 |
| 28 | phi-3-small 7.4B | 69.33 | self-reported· optimizedT1 | 2024-04-23 |
| 29 | Qwen2.5-Coder (7B) | 69.07 | self-reported· optimizedT1 | 2024-09-18 |
| 30 | phi-3-mini 3.8B | 68.93 | self-reported· optimizedT1 | 2024-04-23 |
| 31 | MPT-7B | 68.53 | self-reported· optimizedT1 | 2023-05-05 |
| 32 | Yi-9B | 68.53 | unverified· optimizedT1 | 2024-03-01 |
| 33 | LLaMA-7B | 68.27 | unverified· optimizedT1 | 2023-02-24 |
| 34 | Yi 6B | 65.87 | unverified· optimizedT1 | 2023-11-02 |
| 35 | XGen-7B | 65.6 | self-reported· optimizedT1 | 2023-09-07 |
| 36 | open_llama_7b | 62.4 | self-reported· optimizedT1 | 2023-06-07 |
| 37 | INTELLECT-1 | 61.89 | self-reported· optimizedT1 | 2024-11-29 |
| 38 | Gemma 2B | 61.87 | unverified· optimizedT1 | 2024-02-21 |
| 39 | Qwen2.5-Coder-3B | 61.2 | self-reported· optimizedT1 | 2024-09-18 |
| 40 | Baichuan2-13B | 61.07 | self-reported· optimizedT1 | 2023-09-06 |
| 41 | Dolly 2.0-12b | 61.07 | self-reported· optimizedT1 | 2023-04-12 |
| 42 | internlm-7b | 60.8 | self-reported· optimizedT1 | 2023-07-05 |
| 43 | RedPajama-INCITE-7B-Base | 60.4 | self-reported· optimizedT1 | 2023-05-04 |
| 44 | Baichuan 2-7B | 57.33 | self-reported· optimizedT1 | 2023-09-20 |
| 45 | Qwen2.5-Coder (1.5B) | 49.07 | self-reported· optimizedT1 | 2024-09-18 |
| 46 | Cerebras-GPT-13B | 45.87 | self-reported· optimizedT1 | 2023-04-06 |
| 47 | vicuna-13b-v1.1 | 43.73 | self-reported· optimizedT1 | 2023-04-12 |
| 48 | chatglm2-6b | 42.67 | self-reported· optimizedT1 | 2023-06-24 |
| 49 | Phi-2 | 38.13 | self-reported· optimizedT1 | 2023-12-12 |
| 50 | Qwen2.5-Coder-0.5B | 31.2 | self-reported· optimizedT1 | 2024-09-18 |
| 51 | Phi-1.5 | 30.13 | self-reported· optimizedT1 | 2023-09-11 |
| 52 | stablelm-tuned-alpha-7b | 20.93 | self-reported· optimizedT1 | 2023-04-19 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
This benchmark scores 52 models, with a spread of 72.8 points between best and worst — a wide gap that supports genuine separation in capabilities rather than random noise.
2 cited facts
With an Integrity Index of 67 (grade C) and a rank of 43rd among 51 benchmarks, the benchmark's health as a discriminator of frontier models is compromised most by contamination, which narrows how much to trust top scores.
5 cited facts
The benchmark is not saturated, as the top score of 93.73 leaves a headroom of 6.27 points, and only one model clusters near the top, indicating that the ranking retains genuine discrimination at the peak.
4 cited facts
This benchmark uses a mixed harness with low comparability, meaning scores from different harnesses should not be compared directly. The test set is public, so high performance requires scrutiny for possible contamination — notably, the data has been publicly available for an extended period and is widely included in pretraining corpora, leading to it being treated as contaminated. The harness itself is sensitive to scoring format, with length-normalized log-likelihood and generative variants causing significant swings, and official test labels are behind a leaderboard so validation is typically reported.
4 cited facts
Sentence completion is an indirect probe: predicting the likely next event is related to commonsense but not the same as applying it, so keep the score scoped to the format. It was built using adversarial filtering over video and procedural-text contexts, creating four-choice multiple-choice items. The proxy is the vulnerability: rejecting adversarially filtered distractors is a format skill adjacent to commonsense reasoning, and the score cannot separate them. A key caveat is that adversarial filtering can produce artifacts such as ungrammatical endings or ambiguous items, so the benchmark's meaningful ceiling is below perfect accuracy and does not capture all aspects of commonsense.
4 cited facts
HellaSwag (HellaSwag (Harder Endings, Longer contexts, and Low-shot Activities for Situations With Adversarial Generations)): A 4-way multiple-choice commonsense sentence-completion benchmark: given a context, pick the most plausible continuation, where wrong endings are machine-generated via Adversarial Filtering. Scored as accuracy and effectively saturated for frontier models.
GPT-4 (Mar 2023) leads HellaSwag at 93.73. The full leaderboard above lists every recorded measurement, not just the headline number.
52 models have recorded scores on HellaSwag, spanning a score spread of 72.8.
tensor.news grades HellaSwag C for integrity (score 67/100), ranking #52 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — HellaSwag still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is public, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Only partly — scores come from a mix of harnesses and protocols, so a head-to-head on HellaSwag is not fully apples-to-apples.
Every HellaSwag measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites arxiv.org/abs/1905.07830.