Editorial
Analysis
Analysis from the evidence record — every number cited to the public dataset.
- Deep dive
GPQA diamond still sorts frontier models — but its public test set makes every score an upper bound
The benchmark's discrimination is intact and 7.2 points of headroom remain, yet a contamination subscore of 5.7/100 means scores above 90% mix genuine reasoning with exposure advantage.
- Deep dive
Gemini 3.1 Pro leads SimpleQA Verified at 77.3% — and 22.7 points of headroom remain
One of the least saturated benchmarks in the index, with a perfect contamination score — but every number on its leaderboard comes from a single evaluator, and that is the limit of what it can prove.
- Brief
Only 29% of AI model performance claims are independently verified
Of 1,943 scores in the public record, 573 are independently reproduced; the pricing picture is worse — more published prices are disputed than corroborated, across a 489-fold cost range.