SimpleQA Verified
Integrity rank #4 of 61 · 68 models scored · top score 77.3 · Gemini 3.1 Pro
unknown
Strongest on discrimination, weakest on saturation headroom. Caveat: the public test set does not rule out training-data contamination.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | Gemini 3.1 Pro | 77.3 | reproducedT1 | 2026-02-19 |
| 2 | Gemini 3 Pro | 72.9 | reproducedT1 | 2025-11-18 |
| 3 | GPT-5.6 Sol | 71.6 | reproducedT1 | 2026-07-09 |
| 4 | Gemini 3.7 Flash | 71.2 | reproducedT1 | 2026-08-13 |
| 5 | Gemini 3.6 Flash | 68.7 | reproducedT1 | 2026-07-21 |
| 6 | Gemini 3.5 Flash | 68.4 | reproducedT1 | 2026-05-19 |
| 7 | Claude Fable 5 | 68.3 | reproducedT1 | 2026-06-09 |
| 8 | Qwen3-Max | 67.47 | reproducedT1 | 2025-09-05 |
| 9 | Gemini 3 Flash | 67.4 | reproducedT1 | 2025-12-17 |
| 10 | Muse Spark | 66.3 | reproducedT1 | 2026-04-08 |
| 11 | GPT-5.5 Pro | 64.5 | reproducedT1 | 2026-04-23 |
| 12 | GPT-5.5 | 63.1 | reproducedT1 | 2026-04-23 |
| 13 | Qwen3.7-Max | 58.52 | reproducedT1 | 2026-05-19 |
| 14 | DeepSeek-V4-Pro | 57 | reproducedT1 | 2026-04-24 |
| 15 | Qwen 3.6 Max (Preview) | 56.93 | reproducedT1 | 2026-04-20 |
| 16 | Claude Opus 5 | 56.7 | reproducedT1 | 2026-07-24 |
| 17 | Gemini 2.5 Pro (Jun 2025) | 56 | reproducedT1 | 2025-06-05 |
| 18 | Grok 4.6 | 54.4 | reproducedT1 | 2026-08-12 |
| 19 | Grok 4.5 | 53.5 | reproducedT1 | 2026-07-08 |
| 20 | o3 | 53 | reproducedT1 | 2024-12-20 |
| 21 | Claude Opus 4.7 | 50.6 | reproducedT1 | 2026-04-16 |
| 22 | GPT-5 | 50.6 | reproducedT1 | 2025-08-07 |
| 23 | Qwen3-235B-A22B-Thinking (Jul 2025) | 50.1 | reproducedT1 | 2025-07-25 |
| 24 | Qwen 3.6 Plus | 49.1 | reproducedT1 | 2026-04-01 |
| 25 | GPT-5.1 | 48.9 | reproducedT1 | 2025-11-13 |
| 26 | Grok 4 | 47.9 | reproducedT1 | 2025-07-09 |
| 27 | GPT-5.4 Pro | 47.8 | reproducedT1 | 2026-03-05 |
| 28 | Claude Opus 4.6 | 46.49 | reproducedT1 | 2026-02-05 |
| 29 | Qwen 3.8 Max | 46.29 | reproducedT1 | 2026-07-19 |
| 30 | GPT-5.4 | 44.83 | reproducedT1 | 2026-03-05 |
| 31 | GPT-5.6 Terra | 43.1 | reproducedT1 | 2026-07-09 |
| 32 | Kimi K3 | 42.7 | reproducedT1 | 2026-07-16 |
| 33 | Claude Opus 4.5 | 41.8 | reproducedT1 | 2025-11-24 |
| 34 | GPT-5.6 Luna | 41.7 | reproducedT1 | 2026-07-09 |
| 35 | Inkling | 40.2 | reproducedT1 | 2026-07-15 |
| 36 | Claude Opus 4.8 | 39.5 | reproducedT1 | 2026-05-28 |
| 37 | Kimi K2.7 Code | 39.2 | reproducedT1 | 2026-06-12 |
| 38 | GPT-5.2 | 38.9 | reproducedT1 | 2025-12-11 |
| 39 | Kimi K2.6 | 38.7 | reproducedT1 | 2026-04-20 |
| 40 | GLM-5.2 | 38.1 | reproducedT1 | 2026-06-16 |
| 41 | Grok 4.3 Beta | 38 | reproducedT1 | 2026-04-17 |
| 42 | GPT-5.5 Instant | 38 | reproducedT1 | 2026-05-05 |
| 43 | GLM-5.1 | 37.29 | reproducedT1 | 2026-04-07 |
| 44 | Grok 4.20 | 35.1 | reproducedT1 | 2026-02-17 |
| 45 | Claude Opus 4.1 | 34.8 | reproducedT1 | 2025-08-05 |
| 46 | DeepSeek V4 Flash 0731 | 34.67 | reproducedT1 | 2026-07-31 |
| 47 | Kimi K2.5 | 33.9 | reproducedT1 | 2026-02-02 |
| 48 | Kimi K2 Thinking | 31.6 | reproducedT1 | 2025-11-06 |
| 49 | GLM-4.7 | 31.5 | reproducedT1 | 2025-12-22 |
| 50 | Claude Sonnet 4.6 | 29 | reproducedT1 | 2026-02-17 |
| 51 | GPT-5.4 Mini | 28.6 | reproducedT1 | 2026-03-17 |
| 52 | DeepSeek-V3.2 | 27.5 | reproducedT1 | 2025-12-01 |
| 53 | DeepSeek-R1 (May 2025) | 27.4 | reproducedT1 | 2025-05-28 |
| 54 | Qwen 3.5 Plus (hosted 397B-A17B) | 26 | reproducedT1 | 2026-02-16 |
| 55 | Claude Sonnet 5 | 25 | reproducedT1 | 2026-06-30 |
| 56 | o4-mini | 23.9 | reproducedT1 | 2025-04-16 |
| 57 | Claude Sonnet 4.5 | 23.6 | reproducedT1 | 2025-09-29 |
| 58 | Qwen 3.6 Flash | 21.16 | reproducedT1 | 2026-04-26 |
| 59 | Grok-3 mini | 21.1 | reproducedT1 | 2025-02-19 |
| 60 | GPT-5 mini | 21 | reproducedT1 | 2025-08-07 |
| 61 | Qwen 3.5 Flash (hosted 35B-A3B) | 19.76 | reproducedT1 | 2026-02-25 |
| 62 | Inkling-Small | 19.5 | reproducedT1 | 2026-07-30 |
| 63 | gpt-oss-120b | 13.9 | reproducedT1 | 2025-08-05 |
| 64 | GPT-5 nano | 12.2 | reproducedT1 | 2025-08-07 |
| 65 | GPT-5.4 Nano | 12 | reproducedT1 | 2026-03-17 |
| 66 | Gemma 4 31B IT | 9.56 | reproducedT1 | 2026-04-02 |
| 67 | Claude 3.5 Haiku | 6.7 | reproducedT1 | 2024-10-22 |
| 68 | Claude Haiku 4.5 | 5.9 | reproducedT1 | 2025-10-15 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
This benchmark ranks 68 models, with a 71.4-point spread between the best and worst scores; that wide spread suggests the observed differences reflect a genuine capability gap rather than noise.
2 cited facts
The benchmark's Benchmark Integrity Index is 100 (grade A), ranking 4th among 61 benchmarks. Its weakest integrity component is saturation, meaning the score ceiling is crowded. This integrity score measures the benchmark's health as a discriminator of frontier models under a disclosed harness, rather than any model's capability.
6 cited facts
The benchmark is not saturated: the leading score is 77.3, leaving 22.7 points of headroom, and only one model is within a couple of points of the top. This means the ranking still offers genuine room to separate the field at the top, rather than treating small differences among leading models as meaningful capability gaps.
5 cited facts
Scores are generated under a consistent harness, so direct head-to-head comparison is appropriate within this disclosed setup. Test set privacy is unknown, so the risk of contamination-related inflation cannot be assessed; a public set would warrant extra scrutiny, whereas a held-out set would not. Contamination history is unknown, so scores should be read strictly as task performance under the disclosed harness rather than as deployed capability.
3 cited facts
This benchmark's intended measurement target is not documented in the available facts. Its construction process is not documented, so how the test was built cannot be described from these facts. The assumptions underlying this benchmark are not documented, which limits how its results can be interpreted. The sharpest caveat is that this benchmark's limitations are not documented, so readers cannot know where it is most likely to mislead.
4 cited facts
SimpleQA Verified: SimpleQA Verified as reported in Epoch AI's Capabilities Index CSV.
Gemini 3.1 Pro leads SimpleQA Verified at 77.3 — independently reproduced. The full leaderboard above lists every recorded measurement, not just the headline number.
68 models have recorded scores on SimpleQA Verified, spanning a score spread of 71.4.
tensor.news grades SimpleQA Verified A for integrity (score 100/100), ranking #4 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — SimpleQA Verified still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is unknown, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Scores are reported under a consistent harness, so comparisons on SimpleQA Verified are reasonably apples-to-apples.
Every SimpleQA Verified measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites epoch.ai/data/eci_benchmarks.csv.