LAnguage Modeling Broadened to Account for Discourse Aspects
Integrity rank #61 of 61 · 20 models scored · top score 79.8 · Falcon-180B
Long-range, discourse-dependent word prediction
Strongest on saturation headroom, weakest on contamination resistance. Caveat: scores come from an inconsistent mix of harnesses, so head-to-head comparisons are not apples-to-apples.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | Falcon-180B | 79.8 | unverified· optimizedT1 | 2023-09-06 |
| 2 | Llama 2-70B | 78.9 | self-reported· optimizedT1 | 2023-07-18 |
| 3 | LLaMA-65B | 77.7 | self-reported· optimizedT1 | 2023-02-24 |
| 4 | Falcon-40B | 77.3 | unverified· optimizedT1 | 2023-03-15 |
| 5 | LLaMA-33B | 77.2 | self-reported· optimizedT1 | 2023-02-27 |
| 6 | Llama 2-13B | 76.5 | self-reported· optimizedT1 | 2023-07-18 |
| 7 | LLaMA-13B | 75.2 | self-reported· optimizedT1 | 2023-02-27 |
| 8 | Falcon-7B | 74.9 | unverified· optimizedT1 | 2023-04-24 |
| 9 | Baichuan2-13B | 74 | self-reported· optimizedT1 | 2023-09-06 |
| 10 | Llama 2-7B | 73.3 | self-reported· optimizedT1 | 2023-07-18 |
| 11 | LLaMA-7B | 73.3 | self-reported· optimizedT1 | 2023-02-24 |
| 12 | Baichuan 2-7B | 73.3 | self-reported· optimizedT1 | 2023-09-20 |
| 13 | internlm-20b | 71.8 | self-reported· optimizedT1 | 2023-09-18 |
| 14 | Stable Beluga 2 | 71.3 | self-reported· optimizedT1 | 2023-07-20 |
| 15 | Qwen-14B | 71.1 | self-reported· optimizedT1 | 2023-09-24 |
| 16 | MPT-7B | 70 | self-reported· optimizedT1 | 2023-05-05 |
| 17 | Qwen-7B | 67.9 | self-reported· optimizedT1 | 2023-09-28 |
| 18 | internlm-7b | 67 | self-reported· optimizedT1 | 2023-07-05 |
| 19 | Qwen-1_8B | 58.4 | self-reported· optimizedT1 | 2023-11-30 |
| 20 | chatglm2-6b | 54.3 | self-reported· optimizedT1 | 2023-06-24 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
Only 25.5 points separate best from worst among the 20 models here, so the field is bunched and close ranks are weak evidence of a real capability gap. This wide spread suggests that rank differences reflect a genuine capability gap rather than measurement noise.
3 cited facts
The benchmark's Benchmark Integrity Index is 54 (grade D), ranking 51st out of 51 benchmarks, reflecting its health as a discriminator of frontier models under a disclosed harness. Its weakest component is contamination, which narrows how much to trust top scores.
5 cited facts
Ceiling pressure is not the story here: 20.2 points sit above the 79.8 leader, so gaps between leading models are measured distance and the ranking still has room to change. Two models cluster near the top, within a couple of points of the score, meaning that there is still genuine room to separate the field at the top.
6 cited facts
Task performance on this benchmark under a disclosed harness is not directly comparable across different harnesses because the harness is mixed; the test set is public, so high scores deserve increased scrutiny for possible contamination. The mixed harness compounds a known contamination history and the sensitivity of the metric to tokenization and target matching details, further complicating score interpretation.
5 cited facts
Long-range word prediction is a narrow probe of whether distant context is actually used; scores speak to context sensitivity, not to comprehension in any broader sense. It was constructed from thousands of passages drawn from BookCorpus novels, filtered so that humans can guess the target word given the whole passage but not from the last sentence alone. The benchmark assumes that predicting the final word requires integrating broad discourse context, so accuracy serves as a proxy for long-range language modeling. A key caveat is that multiple valid last words are sometimes possible, which means the benchmark may penalize models that choose a different but plausible word.
4 cited facts
LAMBADA (LAnguage Modeling Broadened to Account for Discourse Aspects): A language-modeling benchmark that asks a model to predict the final word of a passage whose last sentence is only resolvable using the full preceding context; scored as last-word accuracy (and sometimes perplexity).
Falcon-180B leads LAMBADA at 79.8. The full leaderboard above lists every recorded measurement, not just the headline number.
20 models have recorded scores on LAMBADA, spanning a score spread of 25.5.
tensor.news grades LAMBADA D for integrity (score 54/100), ranking #61 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — LAMBADA still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is public, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Only partly — scores come from a mix of harnesses and protocols, so a head-to-head on LAMBADA is not fully apples-to-apples.
Every LAMBADA measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites arxiv.org/abs/1606.06031.