The evidence layer for AI model claims
A score is f(model, harness, protocol), not f(model). We evaluate the evaluations — grading each benchmark on discrimination, saturation, contamination, and harness comparability — so you know which numbers still mean something.
61 benchmarks graded, most trustworthy first. Every input is source-backed.
| # | Grade | Benchmark | Signals | Top score |
|---|---|---|---|---|
| 1 | A100 | ARC-AGI-2 | consistent harnessheld-out test set | GPT-5.6 Sol 92.5 · 71 models |
| 2 | A100 | Chess Puzzles | consistent harnessheld-out test set | GPT-5.5 Pro 62.12 · 81 models |
| 3 | A100 | HLE | consistent harnesspublic test set | Gemini 3.1 Pro 43.74 · 37 models |
| 4 | A100 | SimpleQA Verified | consistent harnesspublic test set | Gemini 3.1 Pro 77.3 · 68 models |
| 5 | A99 | APEX-Agents | consistent harnesspublic test set | Claude Fable 5 45 · 47 models |
| 6 | A99 | EBR-bench | consistent harnesspublic test set | Claude Opus 5 50 · 18 models |
| 7 | A99 | FrontierCode | consistent harnesspublic test set | Claude Fable 5 53.5 · 25 models |
| 8 | A99 | FrontierMath-2025-02-28-Private | consistent harnessheld-out test set | Muse Spark 68.42 · 45 models |
| 9 | A99 | FrontierMath-Tier-4-v2-Private | consistent harnessheld-out test set | Claude Fable 5 87.8 · 48 models |
| 10 | A99 | FrontierMath-Tiers-1-3-v2-Private | consistent harnessheld-out test set | GPT-5.6 Sol 89.12 · 48 models |
| 11 | A99 | Mystery Game Puzzles | consistent harnesspublic test set | Claude Opus 5 54.84 · 35 models |
| 12 | A99 | ProofBench | consistent harnesspublic test set | Claude Opus 5 99 · 51 models |
| 13 | A98 | CadEval | consistent harnesspublic test set | o3 74 · 14 models |
| 14 | A98 | DeepSWE | consistent harnesspublic test set | Claude Opus 5 73.65 · 21 models |
| 15 | A98 | GDPval | consistent harnessheld-out test set | GPT-5.2 49.7 · 11 models |
| 16 | A98 | LiveCodeBench | consistent harnesspublic test set | DeepSeek-R1 65.9 · 8 models |
| 17 | A98 | SimpleBench | consistent harnessheld-out test set | Claude Fable 5 78.28 · 71 models |
| 18 | A97 | AIME | consistent harnesspublic test set | DeepSeek-R1 79.8 · 8 models |
| 19 | A97 | MirrorCode | consistent harnesspublic test set | Claude Fable 5 63.89 · 6 models |
| 20 | A96 | Codeforces rating | consistent harnesspublic test set | DeepSeek-R1 2029 · 8 models |
| 21 | A96 | Surface Evolver Bench | consistent harnesspublic test setsaturated | Kimi K3 95 · 19 models |
| 22 | A95 | PostTrainBench | consistent harnesspublic test set | Claude Fable 5 41.79 · 25 models |
| 23 | A94 | CritPt | consistent harnessheld-out test set | GPT-5.6 Sol 32.3 · 76 models |
| 24 | A94 | ExploitBench | consistent harnesspublic test set | GPT-5.5 47.4 · 8 models |
| 25 | A93 | FrontierMath-Tier-4-2025-07-01-Private | consistent harnessheld-out test set | Gemini 3 Pro 31.25 · 27 models |
| 26 | A91 | METR Time Horizons | mixed harness — not comparablepublic test set | Claude Opus 4.6 78.86 · 37 models |
| 27 | A91 | The Agent Company | consistent harnesspublic test set | DeepSeek-V3.2-Exp 42.9 · 14 models |
| 28 | A89 | Balrog | consistent harnesspublic test set | Gemini 3 Pro 58.1 · 24 models |
| 29 | A89 | GSO-Bench | consistent harnesspublic test set | Claude Opus 4.8 47.06 · 24 models |
| 30 | A88 | GeoBench | consistent harnesspublic test set | Gemini 3 Flash 88 · 26 models |
| 31 | A87 | Aider polyglot | consistent harnesspublic test set | GPT-5 88 · 43 models |
| 32 | A87 | WeirdML | consistent harnessheld-out test set | Claude Fable 5 91.94 · 108 models |
| 33 | A86 | Fiction.LiveBench | consistent harnesspublic test set | o3-pro 97.2 · 43 models |
| 34 | A86 | OTIS Mock AIME 2024-2025 | consistent harnesspublic test setsaturated | Claude Fable 5 100 · 136 models |
| 35 | A86 | VPCT | consistent harnesspublic test set | Gemini 3 Pro 86.5 · 26 models |
| 36 | B84 | Terminal Bench | consistent harnesspublic test set | GPT-5.5 84.7 · 36 models |
| 37 | B83 | SWE-Bench verified | consistent harnesspublic test set | Claude Opus 4.7 83.47 · 32 models |
| 38 | B82 | CL-bench | consistent harnesspublic test set | GPT-5.4 27.9 · 19 models |
| 39 | B82 | OSWorld 2.0 | consistent harnesspublic test set | Claude Opus 4.8 20.6 · 7 models |
| 40 | B81 | CL-bench Life | consistent harnesspublic test set | GPT-5.5 22.2 · 13 models |
| 41 | B81 | Remote Labor Index | consistent harnesspublic test set | Claude Fable 5 16.1 · 10 models |
| 42 | B80 | OSWorld | consistent harnesspublic test set | Claude Sonnet 4.6 72.1 · 8 models |
| 43 | B79 | Cybench | mixed harness — not comparablepublic test set | Claude Opus 4.6 93 · 19 models |
| 44 | B77 | DeepResearch Bench | consistent harnesspublic test set | Claude Opus 4.6 55.31 · 22 models |
| 45 | B76 | Lech Mazur Writing | consistent harnesspublic test set | GPT-5 86 · 41 models |
| 46 | B76 | MATH-500 | consistent harnesspublic test set | DeepSeek-R1 97.3 · 8 models |
| 47 | B73 | ARC-AGI | consistent harnesspublic test setsaturated | Claude Fable 5 98.5 · 72 models |
| 48 | B73 | GPQA diamond | mixed harness — not comparablepublic test set | Gemini 3.7 Flash 93.1 · 153 models |
| 49 | B73 | MATH level 5 | consistent harnesspublic test setsaturated | GPT-5 98.13 · 79 models |
| 50 | B71 | BBH | mixed harness — not comparablepublic test set | Gemini 1.5 Pro (May 2024) 85.6 · 41 models |
| 51 | C68 | GSM8K | mixed harness — not comparablepublic test set | GPT-4 (Mar 2023) 92 · 56 models |
| 52 | C67 | HellaSwag | mixed harness — not comparablepublic test set | GPT-4 (Mar 2023) 93.73 · 52 models |
| 53 | C67 | MMLU | mixed harness — not comparablepublic test set | GPT-4o (Nov 2024) 84.13 · 99 models |
| 54 | C67 | ScienceQA | mixed harness — not comparablepublic test set | GPT-4o (May 2024) 84.67 · 6 models |
| 55 | C66 | ARC AI2 | mixed harness — not comparablepublic test set | DeepSeek-V3 93.73 · 66 models |
| 56 | C63 | TriviaQA | mixed harness — not comparablepublic test set | Llama 2-70B 87.6 · 32 models |
| 57 | C62 | PIQA | mixed harness — not comparablepublic test set | PowerMoE-3b 79.1 · 44 models |
| 58 | C62 | Winogrande | mixed harness — not comparablepublic test set | Llama 3.1-405B 78.4 · 64 models |
| 59 | C61 | OpenBookQA | mixed harness — not comparablepublic test set | phi-3-mini 3.8B 84 · 33 models |
| 60 | C59 | ANLI | consistent harnesspublic test set | phi-3-small 7.4B 37.15 · 9 models |
| 61 | D54 | LAMBADA | mixed harness — not comparablepublic test set | Falcon-180B 79.8 · 20 models |
Integrity = 0.30·discrimination + 0.30·saturation + 0.15·contamination + 0.15·harness + 0.10·age. A high grade means the benchmark still separates models, is not near ceiling, resists training-set contamination, and is scored under a consistent harness. It does not mean the number reflects real-world or agentic performance. That is a separate question we flag but never conflate. Full methodology →