Skip to content

GPQA diamond still sorts frontier models — but its public test set makes every score an upper bound

tensor.news desk · 2026-07-24

GPQA diamond is a multiple-choice benchmark that tests graduate-level reasoning in biology, chemistry, and physics. It was released in November 2023 by researchers at NYU, Cohere, and Anthropic. The explicit design goal was to create questions that are difficult for humans even when they have unrestricted web access, making it a proxy for deep domain expertise rather than surface-level retrieval. The metric is straightforward: 4-option accuracy. That simplicity, combined with a large model count of 122 submissions, has made GPQA diamond a standard reference point for frontier capabilities.

What the benchmark actually measures is narrow but well-defined. Each question demands synthesis of graduate-level concepts; the "Google-proof" framing means that simple search-and-extract strategies fail. In practice, the benchmark evaluates a model's ability to chain facts, reason through counterfactuals, and identify correct answers when all answer choices appear plausible to a non-expert. That design keeps the signal high for advanced reasoning, which is why the discrimination component of its Benchmark Integrity Index stands at 100. The spread between the weakest and strongest models among 122 entries is 92.8 percentage points, confirming that the benchmark continues to separate capability levels cleanly.

Despite its reputation for difficulty, GPQA diamond is not saturated. The saturation flag is false. The clustering ratio is 0.041, meaning top scores are not bunched so tightly that differentiation disappears. True, the best score of 92.8% is close to the ceiling, but 7.2 percentage points of headroom remain. A model with perfect reasoning would still have room to improve, and the current top of the leaderboard shows a gradient: GPT-5.4 Pro at 92.8%, Gemini 3.1 Pro at 92.13%, GPT-5.5 at 92%, and Qwen3.7-Max at 88.8%. Those gaps, though small in absolute terms, are meaningful across the cluster of frontier systems. From a ranking standpoint, the benchmark has not yet exhausted its sorting power.

Where the difficulty breaks down is contamination. The Benchmark Integrity Index assigns a contamination score of 5.7 out of 100, the weakest component of the benchmark's overall B grade (rank 31 of 51). The test set is explicitly public. The contamination window ratio is 0.943, indicating that almost the entire evaluation period overlaps with publicly available data that could have seeped into training corpora. Even though the questions were originally designed to resist web search, the public release of the question set and answers has made it plausible that models have memorized the items, or have been fine-tuned on evaluation data leakage. This contamination risk does not make the benchmark useless, but it shifts the interpretation: a high score reflects a mix of genuine reasoning and exposure advantage rather than pure out-of-distribution capability.

Who leads the leaderboard is clear: GPT-5.4 Pro holds the top spot at 92.8%, followed by Gemini 3.1 Pro (92.13%) and GPT-5.5 (92%). Every score in the top 10 comes from Epoch AI evaluations, and every entry is marked as independently reproduced and optimized. The harness is consistent across all listed results — no mixed evaluation frameworks, no self-reported numbers from model developers. That gives strong internal consistency. However, the evidence for each score is single-source: none of the top results have been cross-checked by a second independent evaluator. This is not a disqualifier, but it limits confidence — if a single evaluator introduces systematic bias, even unintentionally, there is no external validation to detect it. Within the same harness, scores are comparable, but no multi-lab consensus exists as additional assurance.

The age of the benchmark — 977 days at the time of writing — further erodes its integrity. An age score of 66.5 reflects a typical decay in freshness. As models grow stronger and the community invests heavily in optimizing for this specific metric, the benchmark inevitably becomes more of a target than a measurement tool. The combination of a public test set, advanced age, and a near-total contamination window means that any score above 90% should be read with the explicit caveat that it likely overstates the model's true reasoning ability in a decontaminated setting.

The one caveat that matters before quoting GPQA diamond: the test set is public and almost certainly contaminated across the full length of the evaluation window. The contamination subscore of 5.7 out of 100 is not a theoretical concern; it reflects a reality where models have had ample opportunity to ingest the exact questions and answers. Consequently, the numeric accuracy values are not pure measures of reasoning. They are best understood as upper-bound estimates of what a model can do when it may have seen the test material during training or fine-tuning. For questions of genuine out-of-distribution reasoning — evaluating readiness for novel scientific problem-solving, say — GPQA diamond alone is insufficient evidence, and needs pairing with private or dynamic benchmarks that control for leakage.

Verdict: GPQA diamond remains the right tool for relative ordering of frontier models within a single, consistent evaluation harness. Epoch AI's uniform harness means the reported values rank models fairly against one another, and the remaining 7.2 points of headroom mean the leaderboard is not flat. It is the wrong tool as a standalone, absolute indicator of graduate-level reasoning or PhD-level expertise: the public test set and severe contamination vulnerability strip its scores of the epistemic safety a high-stakes claim requires. Treat them as an intermediate filter — useful, but never final.