WinoGrande: An Adversarial Winograd Schema Challenge at Scale
Integrity rank #58 of 61 · 64 models scored · top score 78.4 · Llama 3.1-405B
Commonsense pronoun/coreference resolution
Strongest on discrimination, weakest on contamination resistance. Caveat: scores come from an inconsistent mix of harnesses, so head-to-head comparisons are not apples-to-apples.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | Llama 3.1-405B | 78.4 | unverified· optimizedT1 | 2024-07-23 |
| 2 | Claude 3 Opus | 77 | unverified· optimizedT1 | 2024-03-04 |
| 3 | GPT-4 (Mar 2023) | 75 | self-reported· optimizedT1 | 2023-03-15 |
| 4 | Falcon-180B | 74.2 | unverified· optimizedT1 | 2023-09-06 |
| 5 | DeepSeek-V2 (MoE-236B, May 2024) | 72.6 | self-reported· optimizedT1 | 2024-05-07 |
| 6 | DeepSeek-V3 | 70.4 | self-reported· optimizedT1 | 2024-12-24 |
| 7 | Llama 3-70B | 67 | unverified· optimizedT1 | 2024-04-18 |
| 8 | PaLM 2-L | 66 | self-reported· optimizedT1 | 2023-05-17 |
| 9 | PowerMoE-3b | 65 | unverifiedT2 | — |
| 10 | Qwen2.5-72B | 64.6 | self-reported· optimizedT1 | 2024-09-19 |
| 11 | GPT-3.5 Turbo (Jun 2023) | 63.2 | unverified· optimizedT1 | 2023-06-13 |
| 12 | phi-3-medium 14B | 63 | self-reported· optimizedT1 | 2024-04-23 |
| 13 | phi-3-small 7.4B | 63 | self-reported· optimizedT1 | 2024-04-23 |
| 14 | Qwen2.5-Coder-32B | 61.6 | self-reported· optimizedT1 | 2024-09-18 |
| 15 | Llama 2-70B | 60.4 | unverified· optimizedT1 | 2023-07-18 |
| 16 | PaLM 2-M | 58.4 | self-reported· optimizedT1 | 2023-05-17 |
| 17 | Gemma 7B | 58 | unverified· optimizedT1 | 2024-02-21 |
| 18 | Falcon 2 11B | 56.6 | self-reported· optimizedT1 | 2024-05-09 |
| 19 | Nemotron-4 15B | 56 | self-reported· optimizedT1 | 2024-02-27 |
| 20 | PaLM 2-S | 55.8 | self-reported· optimizedT1 | 2023-05-17 |
| 21 | Mixtral 8x7B | 54.4 | unverified· optimizedT1 | 2023-12-11 |
| 22 | LLaMA-65B | 54 | unverified· optimizedT1 | 2023-02-24 |
| 23 | Falcon-40B | 53.8 | unverified· optimizedT1 | 2023-03-15 |
| 24 | Qwen2.5-Coder-14B | 53.6 | self-reported· optimizedT1 | 2024-09-18 |
| 25 | Llama 2-34B | 53.4 | unverified· optimizedT1 | 2023-07-18 |
| 26 | LLaMA-33B | 52 | unverified· optimizedT1 | 2023-02-27 |
| 27 | Llama 3-8B | 51.4 | unverified· optimizedT1 | 2024-04-18 |
| 28 | Mistral 7B v0.1 | 50.6 | unverified· optimizedT1 | 2023-10-10 |
| 29 | Claude 3 Sonnet | 50.2 | unverified· optimizedT1 | 2024-03-04 |
| 30 | Claude 3 Haiku | 48.4 | unverified· optimizedT1 | 2024-03-04 |
| 31 | Phi-1.5 | 46.8 | self-reported· optimizedT1 | 2023-09-11 |
| 32 | LLaMA-13B | 46 | unverified· optimizedT1 | 2023-02-27 |
| 33 | Yi-9B | 46 | unverified· optimizedT1 | 2024-03-01 |
| 34 | Qwen2.5-Coder (7B) | 45.8 | self-reported· optimizedT1 | 2024-09-18 |
| 35 | DeepSeek-Coder-V2-Lite-Base | 45.8 | self-reported· optimizedT1 | 2024-06-13 |
| 36 | Llama 2-13B | 45.6 | unverified· optimizedT1 | 2023-07-18 |
| 37 | Yi 6B | 42.6 | unverified· optimizedT1 | 2023-11-02 |
| 38 | MPT-30B | 42 | unverified· optimizedT1 | 2023-06-22 |
| 39 | phi-3-mini 3.8B | 41.6 | self-reported· optimizedT1 | 2024-04-23 |
| 40 | vicuna-13b-v1.1 | 41.6 | self-reported· optimizedT1 | 2023-04-12 |
| 41 | LLaMA-7B | 40.2 | unverified· optimizedT1 | 2023-02-24 |
| 42 | Llama 2-7B | 38.4 | unverified· optimizedT1 | 2023-07-18 |
| 43 | GPT-3.5 Turbo (Nov 2023) | 37.6 | self-reported· optimizedT1 | 2023-06-13 |
| 44 | MPT-7B | 37.2 | unverified· optimizedT1 | 2023-05-05 |
| 45 | Qwen2.5-Coder-3B | 34.8 | self-reported· optimizedT1 | 2024-09-18 |
| 46 | Falcon-7B | 34.4 | unverified· optimizedT1 | 2023-04-24 |
| 47 | open_llama_7b | 34 | self-reported· optimizedT1 | 2023-06-07 |
| 48 | INTELLECT-1 | 31.64 | self-reported· optimizedT1 | 2024-11-29 |
| 49 | Gemma 2B | 30.8 | unverified· optimizedT1 | 2024-02-21 |
| 50 | XGen-7B | 29.8 | self-reported· optimizedT1 | 2023-09-07 |
| 51 | StarCoder 2 15B | 28.6 | self-reported· optimizedT1 | 2024-02-29 |
| 52 | RedPajama-INCITE-7B-Base | 27.6 | self-reported· optimizedT1 | 2023-05-04 |
| 53 | DeepSeek Coder 33B | 24 | self-reported· optimizedT1 | 2024-01-25 |
| 54 | Dolly 2.0-12b | 23.6 | self-reported· optimizedT1 | 2023-04-12 |
| 55 | Cerebras-GPT-13B | 21.6 | self-reported· optimizedT1 | 2023-04-06 |
| 56 | Qwen2.5-Coder (1.5B) | 21.4 | self-reported· optimizedT1 | 2024-09-18 |
| 57 | CodeQwen1.5-7B | 19.6 | self-reported· optimizedT1 | 2024-04-15 |
| 58 | DeepSeek Coder 6.7B | 15.2 | self-reported· optimizedT1 | 2024-01-25 |
| 59 | StarCoder 2 3B | 14.2 | self-reported· optimizedT1 | 2024-02-29 |
| 60 | StarCoder 2 7B | 14.2 | self-reported· optimizedT1 | 2024-02-29 |
| 61 | Qwen2.5-Coder-0.5B | 9.6 | self-reported· optimizedT1 | 2024-09-18 |
| 62 | Phi-2 | 9.4 | self-reported· optimizedT1 | 2023-12-12 |
| 63 | DeepSeek Coder 1.3B | 6.6 | self-reported· optimizedT1 | 2024-01-25 |
| 64 | stablelm-tuned-alpha-7b | 3 | self-reported· optimizedT1 | 2023-04-19 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
This benchmark scores 64 models, with a spread of 75.4 points between the best and worst; this wide spread indicates that rank differences are likely meaningful and not due to noise.
2 cited facts
The benchmark's Benchmark Integrity Index is 62 (grade C), ranking 48th out of 51 benchmarks. Its weakest component is contamination, which narrows how much to trust top scores. This score reflects the benchmark's health as a discriminator of frontier models under a disclosed harness.
5 cited facts
Room to run remains — 21.6 points above the 78.4 top score — so gaps in this table reflect measured distance rather than ceiling compression, and further separation is still possible. Only 2 models cluster near the top, within a couple of points. Because the benchmark is not saturated, there is still genuine room to separate the field at the top, and small differences among the best models should be considered meaningful.
5 cited facts
This benchmark's harness exhibits mixed comparability, meaning scores from different harnesses should not be directly compared. The test set is public, so high task performance under the disclosed harness merits extra scrutiny for possible contamination, especially given that the data has been publicly available since 2019 and is treated as contaminated in common pretraining corpora. Furthermore, the harness is sensitive to implementation details: it uses few-shot binary choice via log-likelihood, and partial-scoring or normalization choices can swing results.
4 cited facts
Adversarial filtering against surface lexical artifacts is the core defense: it makes strong performance harder to attribute to shallow statistics, which is the main reason to read residual errors as genuine commonsense gaps. Pronoun resolution is a narrow window on commonsense; the leap from resolving referents to reasoning in general is the assumption to hold apart from the score itself. The starkest limitation is that despite the adversarial filtering, residual artifacts remain, some items are ambiguous or exhibit social bias, and current models still underperform relative to human judgment on the task.
4 cited facts
Winogrande (WinoGrande: An Adversarial Winograd Schema Challenge at Scale): A large-scale, adversarially-filtered Winograd Schema benchmark: fill the blank in a near-identical twin-sentence pair with the correct of two referents; scored as binary accuracy.
Llama 3.1-405B leads Winogrande at 78.4. The full leaderboard above lists every recorded measurement, not just the headline number.
64 models have recorded scores on Winogrande, spanning a score spread of 75.4.
tensor.news grades Winogrande C for integrity (score 62/100), ranking #58 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — Winogrande still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is public, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Only partly — scores come from a mix of harnesses and protocols, so a head-to-head on Winogrande is not fully apples-to-apples.
Every Winogrande measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites arxiv.org/abs/1907.10641.