Skip to content

Edition·2026-09-01 · Tue

Fed from·Epoch AI·HF Open LLM·OpenAlex·OpenAI evals

The benchmark record for AI models.

Every reported score is tagged by how it was checked — independently reproduced, vendor self-reported, contradicted, or unverified — and every benchmark is graded on the four signals that decide whether a number still means something. We grade the evaluations, not just the models.

Latest measurement·2026-08-13·35 of 61 benchmarks graded A

At a glancesnapshot 2026-08-13
Benchmarks graded
61
35 rated A
Models tracked
390
Claims in ledger
2,467
Independently reproduced
31%
776 of 2,467

How trustworthy the evaluations are

Benchmark integrity

saturated · belief erodes50607080901000.20.40.60.81.0saturation — top score's proximity to ceiling →integrity →WeirdMLOTIS Mock AIME 2024-2025GPQA diamondLAMBADA
ABCDdot size = models scored · 61 benchmarks

Trust holds the top-left; the bottom-right is where scores max out and belief erodes. The best-graded benchmarks:

A
ARC-AGI-2
A
Chess Puzzles
A
HLE
A
SimpleQA Verified
A
APEX-Agents
A
EBR-bench

Top 10 models by rank

Frontier leaderboard

Scores through 2026-08-13Full leaderboard →
#ModelLabLeadsAvg gapEvidence
1Claude Fable 5Anthropic113.8reproduced
2GPT-5.6 SolOpenAI56.0reproduced
3Claude Opus 5Anthropic44.5reproduced
4GPT-5OpenAI425.9reproduced
5GPT-5.5OpenAI410.0reproduced
6GPT-5.5 ProOpenAI24.3reproduced
7Gemini 3 ProGoogle DeepMind316.2reproduced
8Gemini 3.1 ProGoogle DeepMind218.7reproduced
9Claude Opus 4.6Anthropic318.0reproduced
10GPT-4 (Mar 2023)OpenAI232.4reproduced

Leads = benchmarks topped; avg gap = mean points behind SOTA elsewhere. Evidence tags how each score was checked — a score is f(model, harness, protocol), not f(model).

Head-to-head

Popular matchups

Follow the record

Every new score, contradiction, and analysis as it lands — plain Atom, open in any feed reader. No account, no inbox.

Section map

Explore the record