Benchmark explainer
What is GPQA diamond?
Graduate-level bio/chem/physics reasoning, hard even with web access
GPQA Diamond — the hardest 198-question subset of Graduate-level Google-Proof Q&A across biology, chemistry, and physics, authored by PhD-level experts; scored as accuracy (%). PhD-level human experts score ~70% and random guessing ~25%.
How it's scored
- Metric
- MC accuracy (4-option)
- Score ceiling
- 100
- Construction
- Expert-written PhD-validated; Diamond=198 items both experts agreed & non-experts failed
- Human baseline
- PhD experts ~65% (74% excl. errors); non-experts w/web ~34%
How trustworthy
Discrimination
100/100Does it still separate models?
Saturation headroom
93.9/100How far from ceiling / clustered at the top?
Contamination resistance
5.2/100Public vs held-out; training-leak risk.
Harness comparability
50/100Is it apples-to-apples, or mixed / vendor-optimized?
Freshness
64.8/100How old is the benchmark?
Contamination history: Public since 2023 w/ canary but widely circulated; leakage likely for 2025-26 models
Limitations: Small MC set; recall+elimination not open-ended research; format/self-consistency worth a few pts
Who leads GPQA diamond
| Model | Score | Evidence |
|---|---|---|
| Gemini 3.7 Flash | 93.1 | reproduced |
| GPT-5.4 Pro | 92.8 | reproduced |
| Gemini 3.1 Pro | 92.59 | reproduced |
| Gemini 3.6 Flash | 92.17 | reproduced |
| Grok 4.6 | 92 | reproduced |
Frequently asked questions
GPQA diamond (Graduate-Level Google-Proof Q&A (Diamond)): GPQA Diamond — the hardest 198-question subset of Graduate-level Google-Proof Q&A across biology, chemistry, and physics, authored by PhD-level experts; scored as accuracy (%). PhD-level human experts score ~70% and random guessing ~25%.
Gemini 3.7 Flash leads GPQA diamond at 93.1 — independently reproduced. The full leaderboard above lists every recorded measurement, not just the headline number.
153 models have recorded scores on GPQA diamond, spanning a score spread of 93.1.
tensor.news grades GPQA diamond B for integrity (score 73/100), ranking #48 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — GPQA diamond still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is public, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Only partly — scores come from a mix of harnesses and protocols, so a head-to-head on GPQA diamond is not fully apples-to-apples.
Every GPQA diamond measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites arxiv.org/abs/2311.12022.