The AI model evidence desk· updated 2026-06-16
The desk for AI model performance — ranked by verified benchmark evidence, every score tagged by how it was checked.
- Benchmarks graded
- 51
- 26 rated A
- Models tracked
- 308
- Claims in ledger
- 1,943
- Independently reproduced
- 29%
- 573 of 1,943
Top 10 models by rank
Frontier leaderboard
Updated 2026-06-16Full leaderboard →
| # | Model | Lab | Leads | Avg gap | Evidence |
|---|---|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 6 | 3.3 | reproduced |
| 2 | GPT-5.5 Pro | OpenAI | 4 | 3.1 | reproduced |
| 3 | GPT-5.5 | OpenAI | 4 | 6.3 | reproduced |
| 4 | Gemini 3.1 Pro | Google DeepMind | 3 | 14.2 | reproduced |
| 5 | GPT-5 | OpenAI | 3 | 22.9 | reproduced |
| 6 | Gemini 3 Pro | Google DeepMind | 3 | 14.3 | reproduced |
| 7 | Claude Opus 4.7 | Anthropic | 3 | 14.1 | reproduced |
| 8 | GPT-4 (Mar 2023) | OpenAI | 2 | 30.6 | reproduced |
| 9 | Llama 3.1-405B | Meta AI | 2 | 34.9 | reproduced |
| 10 | Claude Opus 4.6 | Anthropic | 2 | 15.1 | reproduced |
Leads = benchmarks topped; avg gap = mean points behind SOTA elsewhere. Evidence tags how each score was checked — a score is f(model, harness, protocol), not f(model).
How trustworthy the evaluations are
Benchmark integrity
Trust holds the top-left; the bottom-right is where scores max out and belief erodes. The best-graded benchmarks:
A | APEX-Agents | consistent harness |
A | Chess Puzzles | consistent harness |
A | HLE | consistent harness |
A | SimpleQA Verified | consistent harness |
A | ARC-AGI-2 | consistent harness |
A | FrontierMath-2025-02-28-Private | consistent harness |
Head-to-head
Popular matchups
Section map