Skip to content

The AI model evidence desk· updated 2026-06-16

The desk for AI model performance — ranked by verified benchmark evidence, every score tagged by how it was checked.

Benchmarks graded
51
26 rated A
Models tracked
308
Claims in ledger
1,943
Independently reproduced
29%
573 of 1,943

Top 10 models by rank

Frontier leaderboard

Updated 2026-06-16Full leaderboard →
#ModelLabLeadsAvg gapEvidence
1Claude Fable 5Anthropic63.3reproduced
2GPT-5.5 ProOpenAI43.1reproduced
3GPT-5.5OpenAI46.3reproduced
4Gemini 3.1 ProGoogle DeepMind314.2reproduced
5GPT-5OpenAI322.9reproduced
6Gemini 3 ProGoogle DeepMind314.3reproduced
7Claude Opus 4.7Anthropic314.1reproduced
8GPT-4 (Mar 2023)OpenAI230.6reproduced
9Llama 3.1-405BMeta AI234.9reproduced
10Claude Opus 4.6Anthropic215.1reproduced

Leads = benchmarks topped; avg gap = mean points behind SOTA elsewhere. Evidence tags how each score was checked — a score is f(model, harness, protocol), not f(model).

How trustworthy the evaluations are

Benchmark integrity

saturated · belief erodes50607080901000.20.40.60.81.0saturation — top score's proximity to ceiling →integrity →OTIS Mock AIME 2024-2025GPQA diamondMMLULAMBADA
ABCDdot size = models scored · 51 benchmarks

Trust holds the top-left; the bottom-right is where scores max out and belief erodes. The best-graded benchmarks:

A
APEX-Agents
A
Chess Puzzles
A
HLE
A
SimpleQA Verified
A
ARC-AGI-2
A
FrontierMath-2025-02-28-Private

Head-to-head

Popular matchups

Section map

Explore the record