Graduate-Level Google-Proof Q&A (Diamond)
Integrity rank #48 of 61 · 153 models scored · top score 93.1 · Gemini 3.7 Flash
Graduate-level bio/chem/physics reasoning, hard even with web access
Strongest on discrimination, weakest on contamination resistance. Caveat: scores come from an inconsistent mix of harnesses, so head-to-head comparisons are not apples-to-apples.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | Gemini 3.7 Flash | 93.1 | reproduced· optimizedT1 | 2026-08-13 |
| 2 | GPT-5.4 Pro | 92.8 | reproduced· optimizedT1 | 2026-03-05 |
| 3 | Gemini 3.1 Pro | 92.59 | reproduced· optimizedT1 | 2026-02-19 |
| 4 | Gemini 3.6 Flash | 92.17 | reproduced· optimizedT1 | 2026-07-21 |
| 5 | Grok 4.6 | 92 | reproduced· optimizedT1 | 2026-08-12 |
| 6 | GPT-5.5 | 92 | reproduced· optimizedT1 | 2026-04-23 |
| 7 | GPT-5.5 Pro | 91.9 | reproduced· optimizedT1 | 2026-04-23 |
| 8 | Claude Opus 5 | 91.84 | reproduced· optimizedT1 | 2026-07-24 |
| 9 | GPT-5.6 Sol | 91.33 | reproduced· optimizedT1 | 2026-07-09 |
| 10 | Grok 4.5 | 91.25 | reproduced· optimizedT1 | 2026-07-08 |
| 11 | GPT-5.6 Terra | 91.08 | reproduced· optimizedT1 | 2026-07-09 |
| 12 | GPT-5.4 | 91.07 | reproduced· optimizedT1 | 2026-03-05 |
| 13 | Kimi K3 | 90.82 | reproduced· optimizedT1 | 2026-07-16 |
| 14 | Gemini 3.5 Flash | 90.4 | reproduced· optimizedT1 | 2026-05-19 |
| 15 | Qwen 3.8 Max | 90.24 | reproduced· optimizedT1 | 2026-07-19 |
| 16 | Gemini 3 Pro | 90.15 | reproduced· optimizedT1 | 2025-11-18 |
| 17 | GLM-5.2 | 89.14 | reproduced· optimizedT1 | 2026-06-16 |
| 18 | GPT-5.6 Luna | 88.8 | reproduced· optimizedT1 | 2026-07-09 |
| 19 | GPT-5.2 | 88.53 | reproduced· optimizedT1 | 2025-12-11 |
| 20 | Claude Opus 4.8 | 88.05 | reproduced· optimizedT1 | 2026-05-28 |
| 21 | DeepSeek V4 Flash 0731 | 88.05 | reproduced· optimizedT1 | 2026-07-31 |
| 22 | DeepSeek-V4-Pro | 87.88 | reproduced· optimizedT1 | 2026-04-24 |
| 23 | MiniMax-M3 | 87.88 | reproduced· optimizedT1 | 2026-06-01 |
| 24 | Qwen3.7-Max | 87.88 | reproduced· optimizedT1 | 2026-05-19 |
| 25 | Kimi K2.6 | 87.71 | reproduced· optimizedT1 | 2026-04-20 |
| 26 | Claude Opus 4.6 | 87.37 | reproduced· optimizedT1 | 2026-02-05 |
| 27 | Claude Sonnet 5 | 87.37 | reproduced· optimizedT1 | 2026-06-30 |
| 28 | Claude Opus 4.7 | 86.87 | reproduced· optimizedT1 | 2026-04-16 |
| 29 | GLM-5.1 | 86.53 | reproduced· optimizedT1 | 2026-04-07 |
| 30 | Muse Spark | 86.4 | reproduced· optimizedT1 | 2026-04-08 |
| 31 | Gemini 3 Flash | 85.86 | reproduced· optimizedT1 | 2025-12-17 |
| 32 | Grok 4.20 | 85.77 | reproduced· optimizedT1 | 2026-02-17 |
| 33 | Grok 4.3 Beta | 85.1 | reproduced· optimizedT1 | 2026-04-17 |
| 34 | Inkling-Small | 84.68 | reproduced· optimizedT1 | 2026-07-30 |
| 35 | Qwen 3.6 Plus | 84.51 | reproduced· optimizedT1 | 2026-04-01 |
| 36 | Inkling | 84.34 | reproduced· optimizedT1 | 2026-07-15 |
| 37 | Kimi K2.7 Code | 83.84 | reproduced· optimizedT1 | 2026-06-12 |
| 38 | Qwen3.7-Plus | 83.84 | reproduced· optimizedT1 | 2026-06-02 |
| 39 | GLM-5 | 83.76 | reproduced· optimizedT1 | 2026-02-11 |
| 40 | GPT-5.1 | 83.5 | reproduced· optimizedT1 | 2025-11-13 |
| 41 | Kimi K2.5 | 83.47 | reproduced· optimizedT1 | 2026-02-02 |
| 42 | Qwen 3.6 Max (Preview) | 83.16 | reproduced· optimizedT1 | 2026-04-20 |
| 43 | Claude Sonnet 4.6 | 83.16 | reproduced· optimizedT1 | 2026-02-17 |
| 44 | Grok 4 | 82.67 | reproduced· optimizedT1 | 2025-07-09 |
| 45 | GPT-5.4 Mini | 82.49 | reproduced· optimizedT1 | 2026-03-17 |
| 46 | GPT-5 | 81.57 | reproduced· optimizedT1 | 2025-08-07 |
| 47 | Claude Opus 4.5 | 81.4 | reproduced· optimizedT1 | 2025-11-24 |
| 48 | Claude Fable 5 | 81.14 | reproduced· optimizedT1 | 2026-06-09 |
| 49 | Gemini 2.5 Pro (Jun 2025) | 80.39 | reproduced· optimizedT1 | 2025-06-05 |
| 50 | Qwen 3.5 Plus (hosted 397B-A17B) | 79.8 | reproduced· optimizedT1 | 2026-02-16 |
| 51 | Qwen 3.6 35B-A3B | 79.8 | reproduced· optimizedT1 | 2026-04-14 |
| 52 | Kimi K2 Thinking | 78.96 | reproduced· optimizedT1 | 2025-11-06 |
| 53 | Gemini 2.5 Pro (Mar 2025) | 78.45 | reproduced· optimizedT1 | 2025-03-25 |
| 54 | DeepSeek-V3.2 | 77.9 | reproduced· optimizedT1 | 2025-12-01 |
| 55 | Gemini 3.5 Flash-Lite | 77.78 | reproduced· optimizedT1 | 2026-07-21 |
| 56 | Qwen 3.6 Flash | 77.78 | reproduced· optimizedT1 | 2026-04-26 |
| 57 | GLM-4.7 | 77.78 | reproduced· optimizedT1 | 2025-12-22 |
| 58 | GPT-5.5 Instant | 76.68 | reproduced· optimizedT1 | 2026-05-05 |
| 59 | Qwen 3.5 Flash (hosted 35B-A3B) | 76.43 | reproduced· optimizedT1 | 2026-02-25 |
| 60 | Claude Sonnet 4.5 | 76.43 | reproduced· optimizedT1 | 2025-09-29 |
| 61 | Gemini 3.1 Flash-Lite | 75.76 | reproduced· optimizedT1 | 2026-03-03 |
| 62 | o3 | 75.76 | reproduced· optimizedT1 | 2024-12-20 |
| 63 | Qwen3-235B-A22B-Thinking (Jul 2025) | 73.4 | reproduced· optimizedT1 | 2025-07-25 |
| 64 | Claude 3.7 Sonnet | 72.98 | reproduced· optimizedT1 | 2025-02-24 |
| 65 | o4-mini | 72.81 | reproduced· optimizedT1 | 2025-04-16 |
| 66 | Claude Sonnet 4 | 72.25 | reproduced· optimizedT1 | 2025-05-22 |
| 67 | GPT-5.4 Nano | 71.29 | reproduced· optimizedT1 | 2026-03-17 |
| 68 | Claude Opus 4.1 | 69.7 | reproduced· optimizedT1 | 2025-08-05 |
| 69 | o3-mini | 69.36 | reproduced· optimizedT1 | 2025-01-31 |
| 70 | o1 | 69.02 | reproduced· optimizedT1 | 2024-12-05 |
| 71 | DeepSeek-R1 (May 2025) | 68.43 | reproduced· optimizedT1 | 2025-05-28 |
| 72 | Grok-3 mini | 68.35 | reproduced· optimizedT1 | 2025-02-19 |
| 73 | Claude Opus 4 | 68.35 | reproduced· optimizedT1 | 2025-05-22 |
| 74 | gpt-oss-120b | 67.68 | reproduced· optimizedT1 | 2025-08-05 |
| 75 | Gemma 4 31B IT | 67.68 | reproduced· optimizedT1 | 2026-04-02 |
| 76 | Grok 3 | 67.68 | reproduced· optimizedT1 | 2025-02-17 |
| 77 | GPT-5 mini | 66.67 | reproduced· optimizedT1 | 2025-08-07 |
| 78 | DeepSeek-R1-Distill-Llama-70B | 65.2 | self-reported· optimizedT1 | 2025-01-22 |
| 79 | Qwen3-Max | 63.47 | reproduced· optimizedT1 | 2025-09-05 |
| 80 | DeepSeek-R1 | 62.29 | reproduced· optimizedT1 | 2025-01-20 |
| 81 | DeepSeek-R1-Distill-Qwen-32B | 62.1 | self-reported· optimizedT1 | 2025-01-22 |
| 82 | Claude Haiku 4.5 | 61.62 | reproduced· optimizedT1 | 2025-10-15 |
| 83 | Qwen3-235B-A22B | 60.94 | reproduced· optimizedT1 | 2025-04-28 |
| 84 | GPT-5 nano | 59.26 | reproduced· optimizedT1 | 2025-08-07 |
| 85 | DeepSeek-R1-Distill-Qwen-14B | 59.1 | self-reported· optimizedT1 | 2025-01-22 |
| 86 | GPT-4.5 | 58.25 | reproduced· optimizedT1 | 2025-02-27 |
| 87 | DeepSeek-V3 (Mar 2025) | 56.82 | reproduced· optimizedT1 | 2025-03-24 |
| 88 | Llama 4 Maverick | 55.98 | reproduced· optimizedT1 | 2025-04-05 |
| 89 | GPT-4.1 | 55.89 | reproduced· optimizedT1 | 2025-04-14 |
| 90 | Gemini 2.5 Pro (May 2025) | 55.56 | reproduced· optimizedT1 | 2025-05-06 |
| 91 | GPT-4.1 mini | 54.46 | reproduced· optimizedT1 | 2025-04-14 |
| 92 | Gemini 2.0 Pro | 54.21 | reproduced· optimizedT1 | 2024-12-11 |
| 93 | Gemini 2.0 Flash (Feb 2025) | 52.19 | reproduced· optimizedT1 | 2024-12-11 |
| 94 | o1-mini | 49.83 | reproduced· optimizedT1 | 2024-09-12 |
| 95 | DeepSeek-R1-Distill-Qwen-7B | 49.1 | self-reported· optimizedT1 | 2025-01-22 |
| 96 | DeepSeek-R1-Distill-Llama-8B | 49 | self-reported· optimizedT1 | 2025-01-22 |
| 97 | Mistral Medium 3 | 46.04 | reproduced· optimizedT1 | 2025-05-07 |
| 98 | Gemini 1.5 Pro (Sept 2024) | 42.97 | reproduced· optimizedT1 | 2024-09-24 |
| 99 | Gemini 2.0 Flash Thinking (Jan 2025) | 42.76 | reproduced· optimizedT1 | 2025-01-21 |
| 100 | DeepSeek-V3 | 42.05 | reproduced· optimizedT1 | 2024-12-24 |
| 101 | Qwen2.5-Max | 41.5 | reproduced· optimizedT1 | 2025-01-25 |
| 102 | Phi-4 | 41.41 | reproduced· optimizedT1 | 2024-12-12 |
| 103 | Magistral Small 1.0 | 41.41 | reproduced· optimizedT1 | 2025-06-10 |
| 104 | Claude 3.5 Sonnet (October 2024) | 40.4 | reproduced· optimizedT1 | 2024-10-22 |
| 105 | Claude 3.5 Sonnet | 38.72 | reproduced· optimizedT1 | 2024-06-20 |
| 106 | Grok-2 (Dec 2024) | 38.38 | reproduced· optimizedT1 | 2024-08-13 |
| 107 | Llama 4 Scout | 35.77 | reproduced· optimizedT1 | 2025-04-05 |
| 108 | Mistral Large 2 (Nov 2024) | 35.1 | reproduced· optimizedT1 | 2024-07-24 |
| 109 | Llama 3.1-405B | 34.55 | reproduced· optimizedT1 | 2024-07-23 |
| 110 | DeepSeek-R1-Distill-Qwen-1.5B | 33.8 | self-reported· optimizedT1 | 2025-01-22 |
| 111 | o1-preview | 33.75 | reproduced· optimizedT1 | 2024-09-12 |
| 112 | GPT-4o (Aug 2024) | 32.28 | reproduced· optimizedT1 | 2024-05-13 |
| 113 | Qwen2.5-72B | 32.2 | reproduced· optimizedT1 | 2024-09-19 |
| 114 | Mistral Large 2 (Jul 2024) | 32.03 | reproduced· optimizedT1 | 2024-07-24 |
| 115 | GPT-4.1 nano | 31.9 | reproduced· optimizedT1 | 2025-04-14 |
| 116 | GPT-4o (May 2024) | 31.86 | reproduced· optimizedT1 | 2024-05-13 |
| 117 | GPT-4o (Nov 2024) | 30.51 | reproduced· optimizedT1 | 2024-05-13 |
| 118 | Mistral Small 3.1 | 29.97 | reproduced· optimizedT1 | 2025-03-17 |
| 119 | Llama 3.3 70B | 29.92 | reproduced· optimizedT1 | 2024-12-06 |
| 120 | Gemini 1.5 Flash (Sep 2024) | 29.76 | reproduced· optimizedT1 | 2024-05-10 |
| 121 | Claude 3 Opus | 29.55 | reproduced· optimizedT1 | 2024-03-04 |
| 122 | GPT-4 Turbo (Apr 2024) | 28.79 | reproduced· optimizedT1 | 2024-04-09 |
| 123 | gpt-oss-20b | 27.95 | reproduced· optimizedT1 | 2025-08-05 |
| 124 | Gemini 1.5 Pro (May 2024) | 27.82 | reproduced· optimizedT1 | 2024-05-14 |
| 125 | Llama 3.1-70B | 25.59 | reproduced· optimizedT1 | 2024-07-23 |
| 126 | Gemma 3 27B | 25.25 | reproduced· optimizedT1 | 2025-03-12 |
| 127 | GPT-4 Turbo (Nov 2023) | 23.15 | reproduced· optimizedT1 | 2023-11-06 |
| 128 | Llama 3.2 90B | 21.38 | reproduced· optimizedT1 | 2024-09-24 |
| 129 | Qwen2-72B | 21.04 | reproduced· optimizedT1 | 2024-06-07 |
| 130 | Claude 3 Sonnet | 20.79 | reproduced· optimizedT1 | 2024-03-04 |
| 131 | Llama 3-70B | 20.75 | reproduced· optimizedT1 | 2024-04-18 |
| 132 | Gemini 1.5 Flash (May 2024) | 20.5 | reproduced· optimizedT1 | 2024-05-10 |
| 133 | Mistral Large | 18.35 | reproduced· optimizedT1 | 2024-02-26 |
| 134 | Claude 3.5 Haiku | 17.51 | reproduced· optimizedT1 | 2024-10-22 |
| 135 | GPT-4o mini | 16.96 | reproduced· optimizedT1 | 2024-07-18 |
| 136 | Gemma 2 27B | 15.32 | reproduced· optimizedT1 | 2024-06-24 |
| 137 | Claude 3 Haiku | 15.07 | reproduced· optimizedT1 | 2024-03-04 |
| 138 | GPT-4 (Mar 2023) | 14.31 | reproduced· optimizedT1 | 2023-03-15 |
| 139 | Claude 2 | 12.88 | reproduced· optimizedT1 | 2023-07-11 |
| 140 | Mixtral 8x22B | 12.08 | reproduced· optimizedT1 | 2024-04-17 |
| 141 | Gemini 1.0 Pro | 11.95 | reproduced· optimizedT1 | 2023-12-06 |
| 142 | Claude 2.1 | 10.61 | reproduced· optimizedT1 | 2023-11-21 |
| 143 | GPT-4 (Jun 2023) | 7.53 | reproduced· optimizedT1 | 2023-06-13 |
| 144 | Mixtral 8x7B | 7.45 | reproduced· optimizedT1 | 2023-12-11 |
| 145 | Mistral NeMo | 6.52 | reproduced· optimizedT1 | 2024-07-18 |
| 146 | GPT-3.5 Turbo (Nov 2023) | 4.04 | reproduced· optimizedT1 | 2023-06-13 |
| 147 | phi-3-medium 14B | 3.45 | reproduced· optimizedT1 | 2024-04-23 |
| 148 | Gemma 2 9B | 3.28 | reproduced· optimizedT1 | 2024-06-24 |
| 149 | GPT-3.5 Turbo (Jan 2024) | 2.9 | reproduced· optimizedT1 | 2023-06-13 |
| 150 | Llama 2-70B | 1.77 | reproduced· optimizedT1 | 2023-07-18 |
| 151 | Llama 3-8B | 1.43 | reproduced· optimizedT1 | 2024-04-18 |
| 152 | Llama 3.1-8B | 1.26 | reproduced· optimizedT1 | 2024-07-23 |
| 153 | Yi-34B | 0 | reproduced· optimizedT1 | 2023-11-02 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
This benchmark scores 153 models, with a 93.1-point spread between the best and worst results. A spread this wide supports real separation between models, so rank differences are unlikely to reflect mere noise.
3 cited facts
The benchmark's Integrity Index is 73, corresponding to a grade of B, and it ranks 48th among 61 benchmarks. This score reflects the benchmark's health as a discriminator of frontier models under a disclosed harness; its weakest integrity component is contamination, which narrows how much to trust top scores.
6 cited facts
This benchmark is not saturated. The top score is 93.1, leaving 6.9 points of headroom to the ceiling, and ten models cluster within a couple of points of the top. Because it is not saturated, the ranking still has genuine room to separate the field at the top, so differences among leading models are meaningful.
5 cited facts
Scores under this benchmark show mixed harness comparability, so results from different harnesses should not be compared head-to-head. Because the test set is public rather than held out, high scores deserve more scrutiny for possible contamination than they would if it were private. This benchmark has a contamination history and is sensitive to harness details such as chain-of-thought and test-time compute, so treat scores only as task performance under a disclosed harness, not as deployed capability.
4 cited facts
This benchmark is designed to measure graduate-level reasoning in biology, chemistry, and physics, with questions that remain difficult even when web access is allowed. To support that measurement goal, the test items were written by experts and validated by PhDs, and the final set includes only cases where experts agreed on the answer and non-experts could not solve them. Its central assumption is that multiple-choice accuracy on these expert-vetted questions tracks graduate-level reasoning, and that the 'Google-proof' design prevents simple lookup from substituting for that reasoning. The sharpest caveat is that, because this benchmark is a small multiple-choice set, it measures recall and elimination rather than open-ended research ability, and response format or self-consistency can shift scores by a small amount, so it can mislead if read as a full measure of research skill.
7 cited facts
Models with more than one recorded measurement on this benchmark — every one shown, not just the headline number.
GPQA diamond (Graduate-Level Google-Proof Q&A (Diamond)): GPQA Diamond — the hardest 198-question subset of Graduate-level Google-Proof Q&A across biology, chemistry, and physics, authored by PhD-level experts; scored as accuracy (%). PhD-level human experts score ~70% and random guessing ~25%.
Gemini 3.7 Flash leads GPQA diamond at 93.1 — independently reproduced. The full leaderboard above lists every recorded measurement, not just the headline number.
153 models have recorded scores on GPQA diamond, spanning a score spread of 93.1.
tensor.news grades GPQA diamond B for integrity (score 73/100), ranking #48 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — GPQA diamond still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is public, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Only partly — scores come from a mix of harnesses and protocols, so a head-to-head on GPQA diamond is not fully apples-to-apples.
Every GPQA diamond measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites arxiv.org/abs/2311.12022.