Edition·2026-09-01 · Tue
Fed from·Epoch AI·HF Open LLM·OpenAlex·OpenAI evals
The benchmark record for AI models.
Every reported score is tagged by how it was checked — independently reproduced, vendor self-reported, contradicted, or unverified — and every benchmark is graded on the four signals that decide whether a number still means something. We grade the evaluations, not just the models.
Latest measurement·2026-08-13·35 of 61 benchmarks graded A
How trustworthy the evaluations are
Benchmark integrity
Trust holds the top-left; the bottom-right is where scores max out and belief erodes. The best-graded benchmarks:
A | ARC-AGI-2 | consistent harness |
A | Chess Puzzles | consistent harness |
A | HLE | consistent harness |
A | SimpleQA Verified | consistent harness |
A | APEX-Agents | consistent harness |
A | EBR-bench | consistent harness |
Top 10 models by rank
Frontier leaderboard
| # | Model | Lab | Leads | Avg gap | Evidence |
|---|---|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 11 | 3.8 | reproduced |
| 2 | GPT-5.6 Sol | OpenAI | 5 | 6.0 | reproduced |
| 3 | Claude Opus 5 | Anthropic | 4 | 4.5 | reproduced |
| 4 | GPT-5 | OpenAI | 4 | 25.9 | reproduced |
| 5 | GPT-5.5 | OpenAI | 4 | 10.0 | reproduced |
| 6 | GPT-5.5 Pro | OpenAI | 2 | 4.3 | reproduced |
| 7 | Gemini 3 Pro | Google DeepMind | 3 | 16.2 | reproduced |
| 8 | Gemini 3.1 Pro | Google DeepMind | 2 | 18.7 | reproduced |
| 9 | Claude Opus 4.6 | Anthropic | 3 | 18.0 | reproduced |
| 10 | GPT-4 (Mar 2023) | OpenAI | 2 | 32.4 | reproduced |
Leads = benchmarks topped; avg gap = mean points behind SOTA elsewhere. Evidence tags how each score was checked — a score is f(model, harness, protocol), not f(model).
Head-to-head
Popular matchups
Follow the record
Every new score, contradiction, and analysis as it lands — plain Atom, open in any feed reader. No account, no inbox.
Section map
Explore the record
- Leaderboard390
ranked by evidence
- Benchmarks61
graded on integrity
- Compare62
head-to-head matchups
- Labs37
developers tracked
- Papers60
source records
- Methodology
how we rank & grade