OpenBookQA (Open Book Question Answering)
Integrity rank #59 of 61 · 33 models scored · top score 84 · phi-3-mini 3.8B
Multi-hop reasoning that applies a science fact plus common knowledge
Strongest on discrimination, weakest on contamination resistance. Caveat: scores come from an inconsistent mix of harnesses, so head-to-head comparisons are not apples-to-apples.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | phi-3-mini 3.8B | 84 | self-reported· optimizedT1 | 2024-04-23 |
| 2 | phi-3-small 7.4B | 84 | self-reported· optimizedT1 | 2024-04-23 |
| 3 | phi-3-medium 14B | 83.2 | self-reported· optimizedT1 | 2024-04-23 |
| 4 | GPT-3.5 Turbo (Nov 2023) | 81.33 | self-reported· optimizedT1 | 2023-06-13 |
| 5 | Mixtral 8x7B | 81.07 | self-reported· optimizedT1 | 2023-12-11 |
| 6 | Llama 3-8B | 76.8 | self-reported· optimizedT1 | 2024-04-18 |
| 7 | Mistral 7B v0.1 | 73.07 | self-reported· optimizedT1 | 2023-10-10 |
| 8 | Gemma 7B | 71.47 | self-reported· optimizedT1 | 2024-02-21 |
| 9 | Phi-2 | 64.8 | self-reported· optimizedT1 | 2023-12-12 |
| 10 | Falcon-180B | 52.27 | unverified· optimizedT1 | 2023-09-06 |
| 11 | LLaMA-65B | 46.93 | unverified· optimizedT1 | 2023-02-24 |
| 12 | Llama 2-70B | 46.93 | unverified· optimizedT1 | 2023-07-18 |
| 13 | Llama 2-7B | 44.8 | unverified· optimizedT1 | 2023-07-18 |
| 14 | LLaMA-33B | 44.8 | unverified· optimizedT1 | 2023-02-27 |
| 15 | Llama 2-34B | 44.27 | unverified· optimizedT1 | 2023-07-18 |
| 16 | PaLM 2-M | 43.2 | self-reported· optimizedT1 | 2023-05-17 |
| 17 | LLaMA-7B | 42.93 | unverified· optimizedT1 | 2023-02-24 |
| 18 | Llama 2-13B | 42.67 | unverified· optimizedT1 | 2023-07-18 |
| 19 | Falcon-40B | 42.13 | unverified· optimizedT1 | 2023-03-15 |
| 20 | LLaMA-13B | 41.87 | unverified· optimizedT1 | 2023-02-27 |
| 21 | PaLM 2-S | 41.6 | self-reported· optimizedT1 | 2023-05-17 |
| 22 | PowerMoE-3b | 41 | unverifiedT2 | — |
| 23 | MPT-30B | 36 | unverified· optimizedT1 | 2023-06-22 |
| 24 | Falcon-7B | 35.47 | unverified· optimizedT1 | 2023-04-24 |
| 25 | MPT-7B | 35.2 | unverified· optimizedT1 | 2023-05-05 |
| 26 | XGen-7B | 20.27 | self-reported· optimizedT1 | 2023-09-07 |
| 27 | RedPajama-INCITE-7B-Base | 20 | self-reported· optimizedT1 | 2023-05-04 |
| 28 | Dolly 2.0-12b | 18.93 | self-reported· optimizedT1 | 2023-04-12 |
| 29 | open_llama_7b | 18.67 | self-reported· optimizedT1 | 2023-06-07 |
| 30 | Phi-1.5 | 16.27 | self-reported· optimizedT1 | 2023-09-11 |
| 31 | Cerebras-GPT-13B | 14.4 | self-reported· optimizedT1 | 2023-04-06 |
| 32 | vicuna-13b-v1.1 | 10.67 | self-reported· optimizedT1 | 2023-04-12 |
| 33 | stablelm-tuned-alpha-7b | 9.87 | self-reported· optimizedT1 | 2023-04-19 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
Stretched over 74.13 points, the 33 models here separate cleanly enough that score differences on this table generally mean something. Such a wide spread suggests that rank differences likely reflect genuine capability gaps rather than noise.
3 cited facts
The Benchmark Integrity Index is 61 (grade C), ranking 49th out of 51 benchmarks, indicating moderate health as a discriminator of frontier models under the disclosed harness, with contamination as its weakest component, which narrows how much to trust top scores.
5 cited facts
The top score is 84.0 leaving 16.0 points of headroom, and three models cluster within a couple of points of the top, so the benchmark is not yet saturated and small differences among top models still reflect genuine capability gaps.
4 cited facts
This benchmark's scores are not directly comparable across harnesses and the test set is public, implying that cross-harness comparisons are invalid and that high scores should be viewed with caution regarding possible contamination. The test set has been public for an extended period and is part of common pretraining corpora, so contamination is a known issue; moreover, the harness uses a 0/few-shot log-likelihood method on a small test set, which reduces score stability and comparability.
5 cited facts
OpenBookQA (OpenBookQA (Open Book Question Answering)): A 4-way multiple-choice elementary-science QA benchmark requiring a provided 'open book' core science fact to be combined with additional common knowledge; scored as accuracy.
phi-3-mini 3.8B leads OpenBookQA at 84 — self-reported by the lab. The full leaderboard above lists every recorded measurement, not just the headline number.
33 models have recorded scores on OpenBookQA, spanning a score spread of 74.13.
tensor.news grades OpenBookQA C for integrity (score 61/100), ranking #59 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — OpenBookQA still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is public, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Only partly — scores come from a mix of harnesses and protocols, so a head-to-head on OpenBookQA is not fully apples-to-apples.
Every OpenBookQA measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites arxiv.org/abs/1809.02789.