Skip to content

Gemini 3.1 Pro leads SimpleQA Verified at 77.3% — and 22.7 points of headroom remain

tensor.news desk · 2026-07-24

What SimpleQA Verified actually measures

SimpleQA Verified is classified as a reasoning benchmark in the Epoch AI Capabilities Index, but its reported metric is straightforward: accuracy, expressed as a percentage. The task asks models to provide short, factual answers to questions with objectively verifiable correct responses. The "Verified" designation indicates that the evaluation set uses an answer key that has undergone additional validation, though the public record does not enumerate the exact question format, domains, or verification methodology. What is known with certainty is that Epoch AI reports a single numeric value — normalized performance rendered as accuracy — for each model, computed under a consistent test harness.

The benchmark's purpose, therefore, is to measure how often a model retrieves or computes a correct fact when asked directly. It does not assess multi-step reasoning, chain-of-thought, or subjective judgment; the task-type label of "reasoning" is an Epoch AI category, not a description of the required cognitive process. The construction details and original authors are absent from the record, so understanding the exact composition of the question set requires consulting the source Epoch data and its linked documentation.

Why it remains a hard benchmark

SimpleQA Verified earns a perfect score of 100 on the Benchmark Integrity Index, ranking fourth of 51 evaluated benchmarks, with an A grade. That score is driven by near-perfect component ratings: discrimination (100), saturation (98.5), contamination (100), harness consistency (100), and age (100). These numbers indicate a test that cleanly separates model capabilities and is free from known data leakage or measurement artifacts.

The benchmark is far from saturated. Despite a saturation component score of 98.5 — reflecting a well-constructed difficulty distribution — the saturation flag is false. The empirical saturation ratio is only 0.02, meaning there is negligible clustering near the ceiling. The top score among 51 evaluated models is 77.3%, leaving 22.7 percentage points of headroom. The spread across all models is 71.4 points, confirming that the questions continue to produce a wide range of outcomes and that no model has come close to solving the full set. No model has averaged above 78% accuracy, and the leaderboard's tenth-place entry scores below 60%. These facts make SimpleQA Verified one of the least saturated benchmarks in the current index.

Contamination risk is scored at the maximum 100, with no indication of test-set leakage. The age component also receives a perfect score. Test-set privacy, however, is explicitly recorded as unknown — the one open question in an otherwise clean integrity profile.

Who leads, and exactly how much to trust that number

The leaderboard, consisting solely of Epoch AI's own independent evaluations, places Gemini 3.1 Pro at the top with 77.3% accuracy (measured 2026-02-19). It is followed by Gemini 3 Pro (72.9%), Gemini 3.5 Flash (68.4%), Claude Fable 5 (68.3%), and Qwen3-Max (67.47%). Every top-10 entry carries the same evidence profile: not self-reported, independently reproduced by Epoch AI, not optimized, and single-source — the cross-check consensus equals the reported value because no second party has run it.

This means every number on this leaderboard originates from a single evaluator using a consistent harness. Epoch AI's methodology ensures that relative comparisons among these models are valid — the same prompt format, scoring script, and environment — but the absolute accuracy values cannot be corroborated by a second source. A model that scores 77.3% here might score differently if another organization replicated the evaluation with slight prompt variations, even on an identical test set. The trust these figures deserve rests entirely on Epoch AI's rigor, which its open reporting supports — but the lack of cross-validation is the key limitation.

The unit of comparison is safe because harness consistency scores 100: no model in this list used a different prompting method or post-processing flag. The measurement dates range from 2025-09 to 2026-06, and model versions may have changed outside those windows — the numbers are time-stamped observations, not standing properties of the model names they carry.

The one caveat that matters

Before quoting a SimpleQA Verified score, check that the benchmark's task actually matches the intended use. The record provides only a top-level task type ("reasoning") and a metric ("accuracy") — it does not describe the breadth of knowledge, the phrasing style, or the exact process behind the "Verified" answer key. An application demanding nuanced factual reasoning, multi-step inference, or narrow-domain knowledge may not be reflected by this benchmark at all. And because no second evaluator has reproduced the scores, they are best treated as a single powerful experiment rather than a consensus measurement.

Verdict: when to use it, and when not to

SimpleQA Verified is the right tool for a clean, unsaturated, contamination-free ranking of models on short-answer factual accuracy. The A-grade integrity profile, the 22.7-point headroom, and the consistent harness make it well suited to tracking progress over time and to head-to-head comparisons within the Epoch AI evaluation ecosystem.

It is the wrong tool where the requirement is multi-faceted reasoning, cross-validated absolute accuracy from multiple independent labs, or coverage of a domain that may fall outside its undocumented question distribution. In those cases, the single-source provenance and the incomplete task description become significant liabilities. For strict fact-recall benchmarking among models released through mid-2026, however, SimpleQA Verified remains one of the hardest and most reliable signals available.