TriviaQA
Integrity rank #56 of 61 · 32 models scored · top score 87.6 · Llama 2-70B
Factual/trivia knowledge recall and retrieval
Strongest on discrimination, weakest on contamination resistance. Caveat: scores come from an inconsistent mix of harnesses, so head-to-head comparisons are not apples-to-apples.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | Llama 2-70B | 87.6 | unverified· optimizedT1 | 2023-07-18 |
| 2 | Claude 2 | 87.5 | self-reported· optimizedT1 | 2023-07-11 |
| 3 | PaLM 2-L | 86.1 | self-reported· optimizedT1 | 2023-05-17 |
| 4 | LLaMA-65B | 86 | unverified· optimizedT1 | 2023-02-24 |
| 5 | GPT-3.5 Turbo (Nov 2023) | 85.8 | self-reported· optimizedT1 | 2023-06-13 |
| 6 | GPT-4 (Jun 2023) | 84.8 | unverified· optimizedT1 | 2023-06-13 |
| 7 | Llama 2-34B | 84.6 | unverified· optimizedT1 | 2023-07-18 |
| 8 | LLaMA-33B | 83.8 | unverified· optimizedT1 | 2023-02-27 |
| 9 | DeepSeek-V3 | 82.9 | self-reported· optimizedT1 | 2024-12-24 |
| 10 | Llama 3.1-405B | 82.7 | self-reported· optimizedT1 | 2024-07-23 |
| 11 | Mixtral 8x7B | 82.2 | self-reported· optimizedT1 | 2023-12-11 |
| 12 | PaLM 2-M | 81.7 | self-reported· optimizedT1 | 2023-05-17 |
| 13 | DeepSeek-V2 (MoE-236B, May 2024) | 80 | self-reported· optimizedT1 | 2024-05-07 |
| 14 | Falcon-40B | 79.9 | unverified· optimizedT1 | 2023-03-15 |
| 15 | Llama 2-13B | 79.6 | unverified· optimizedT1 | 2023-07-18 |
| 16 | Claude Instant | 78.9 | self-reported· optimizedT1 | 2023-08-09 |
| 17 | LLaMA-13B | 77.9 | unverified· optimizedT1 | 2023-02-27 |
| 18 | Mistral 7B v0.1 | 75.2 | unverified· optimizedT1 | 2023-10-10 |
| 19 | PaLM 2-S | 75.2 | self-reported· optimizedT1 | 2023-05-17 |
| 20 | phi-3-medium 14B | 73.9 | self-reported· optimizedT1 | 2024-04-23 |
| 21 | Llama 2-7B | 73.7 | unverified· optimizedT1 | 2023-07-18 |
| 22 | MPT-30B | 73.6 | unverified· optimizedT1 | 2023-06-22 |
| 23 | Gemma 7B | 72.3 | unverified· optimizedT1 | 2024-02-21 |
| 24 | Qwen2.5-72B | 71.9 | self-reported· optimizedT1 | 2024-09-19 |
| 25 | LLaMA-7B | 71 | unverified· optimizedT1 | 2023-02-24 |
| 26 | Llama 3-8B | 67.7 | self-reported· optimizedT1 | 2024-04-18 |
| 27 | Falcon-7B | 64.6 | unverified· optimizedT1 | 2023-04-24 |
| 28 | phi-3-mini 3.8B | 64 | self-reported· optimizedT1 | 2024-04-23 |
| 29 | MPT-7B | 61.6 | unverified· optimizedT1 | 2023-05-05 |
| 30 | phi-3-small 7.4B | 58.1 | self-reported· optimizedT1 | 2024-04-23 |
| 31 | Gemma 2B | 53.2 | unverified· optimizedT1 | 2024-02-21 |
| 32 | Phi-2 | 45.2 | self-reported· optimizedT1 | 2023-12-12 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
Neither bunched nor stretched, the 42.4-point spread across 32 models puts this table in the middle band: big score gaps are meaningful, small ones are not decisive. This wide spread indicates that the scores reflect real performance differences, supporting reliable rank comparisons.
3 cited facts
The benchmark's Integrity Index is 63 (grade C), ranking 46th out of 51 benchmarks. The weakest component is contamination, which narrows trust in top scores.
4 cited facts
Ranking among the leaders here is tighter than it looks: 5 models sit within a couple of points of the 87.6 top score. The 12.4-point headroom keeps the benchmark discriminating overall, but inside the lead pack small differences deserve little weight. Since it is not saturated, small differences among the top models still reflect genuine capability gaps, not noise.
5 cited facts
Scores from this benchmark have mixed comparability and are not directly comparable across harnesses; the test set is public, so high scores warrant increased scrutiny for potential contamination. The benchmark has been public since 2017 and is included in common pretraining corpora, thus it is treated as contaminated. Task performance is sensitive to harness details such as closed-book versus open-book settings, exact match evaluation depending on answer alias lists and normalization, and the effect of few-shot exemplars on output format.
4 cited facts
Recall is the skill most easily inflated by training-set overlap, so high scores here deserve the contamination question held open rather than answered in the model's favor. Sourcing questions from public trivia websites and pairing them with web evidence places the material squarely inside common training corpora; strong results here are hard to separate from memorization. The design assumes that exact-match on trivia answers adequately proxies factual knowledge in closed-book settings or reading and retrieval abilities in open-book settings. The sharpest caveat is that exact-match scoring undervalues correct answers that differ in alias or format, and the distantly-supervised evidence can be noisy or fail to contain the answer, so the metric may mislead about true model capability.
4 cited facts
TriviaQA: A large trivia question-answering benchmark with independently gathered evidence documents; in LLM evaluation it is run closed- or open-book and scored by Exact Match (and F1).
Llama 2-70B leads TriviaQA at 87.6. The full leaderboard above lists every recorded measurement, not just the headline number.
32 models have recorded scores on TriviaQA, spanning a score spread of 42.4.
tensor.news grades TriviaQA C for integrity (score 63/100), ranking #56 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — TriviaQA still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is public, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Only partly — scores come from a mix of harnesses and protocols, so a head-to-head on TriviaQA is not fully apples-to-apples.
Every TriviaQA measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites arxiv.org/abs/1705.03551.