ProofBench
Integrity rank #12 of 61 · 51 models scored · top score 99 · Claude Opus 5
unknown
Strongest on discrimination, weakest on saturation headroom. Caveat: the public test set does not rule out training-data contamination.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | Claude Opus 5 | 99 | unverifiedT1 | 2026-07-24 |
| 2 | Claude Fable 5 | 95 | unverifiedT1 | 2026-06-09 |
| 3 | Kimi K3 | 87 | unverifiedT1 | 2026-07-16 |
| 4 | GPT-5.6 Sol | 83 | unverifiedT1 | 2026-07-09 |
| 5 | Claude Sonnet 5 | 77 | unverifiedT1 | 2026-06-30 |
| 6 | GPT-5.6 Terra | 71 | unverifiedT1 | 2026-07-09 |
| 7 | Claude Opus 4.8 | 69 | unverifiedT1 | 2026-05-28 |
| 8 | GPT-5.6 Luna | 60 | unverifiedT1 | 2026-07-09 |
| 9 | Gemini 3.7 Flash | 58 | unverifiedT1 | 2026-08-13 |
| 10 | Qwen 3.8 Max | 58 | unverifiedT1 | 2026-07-19 |
| 11 | DeepSeek V4 Flash 0731 | 56 | unverifiedT1 | 2026-07-31 |
| 12 | GPT-5.4 | 56 | unverifiedT1 | 2026-03-05 |
| 13 | Claude Opus 4.7 | 54 | unverifiedT1 | 2026-04-16 |
| 14 | Grok 4.6 | 51 | unverifiedT1 | 2026-08-12 |
| 15 | Claude Opus 4.6 | 50 | unverifiedT1 | 2026-02-05 |
| 16 | GPT-5.5 | 50 | unverifiedT1 | 2026-04-23 |
| 17 | Claude Sonnet 4.6 | 45 | unverifiedT1 | 2026-02-17 |
| 18 | Muse Spark 1.1 | 39 | unverifiedT1 | 2026-07-09 |
| 19 | Gemini 3.6 Flash | 36 | unverifiedT1 | 2026-07-21 |
| 20 | Claude Opus 4.5 | 36 | unverifiedT1 | 2025-11-24 |
| 21 | GLM-5.2 | 35 | unverifiedT1 | 2026-06-16 |
| 22 | Gemini 3.5 Flash | 31 | unverifiedT1 | 2026-05-19 |
| 23 | Grok 4.5 | 31 | unverifiedT1 | 2026-07-08 |
| 24 | Gemini 3.1 Pro | 26 | unverifiedT1 | 2026-02-19 |
| 25 | Qwen3.7-Max | 26 | unverifiedT1 | 2026-05-19 |
| 26 | GLM-5.1 | 22.22 | unverifiedT1 | 2026-04-07 |
| 27 | GPT-5.4 Mini | 21 | unverifiedT1 | 2026-03-17 |
| 28 | Gemini 3 Pro | 20 | unverifiedT1 | 2025-11-18 |
| 29 | Claude Sonnet 4.5 | 19 | unverifiedT1 | 2025-09-29 |
| 30 | MiniMax-M3 | 18 | unverifiedT1 | 2026-06-01 |
| 31 | GPT-5 | 18 | unverifiedT1 | 2025-08-07 |
| 32 | Muse Spark | 17 | unverifiedT1 | 2026-04-08 |
| 33 | DeepSeek-V4-Pro | 16 | unverifiedT1 | 2026-04-24 |
| 34 | Kimi K2.6 | 16 | unverifiedT1 | 2026-04-20 |
| 35 | Gemini 3 Flash | 15 | unverifiedT1 | 2025-12-17 |
| 36 | GPT-5.2 | 15 | unverifiedT1 | 2025-12-11 |
| 37 | Grok 4.20 | 14 | unverifiedT1 | 2026-02-17 |
| 38 | Gemini 3.5 Flash-Lite | 13 | unverifiedT1 | 2026-07-21 |
| 39 | GPT-5 nano | 12 | unverifiedT1 | 2025-08-07 |
| 40 | Grok 4.3 Beta | 11 | unverifiedT1 | 2026-04-17 |
| 41 | Mistral Medium 3.5 | 9 | unverifiedT1 | 2026-04-29 |
| 42 | GPT-5 mini | 9 | unverifiedT1 | 2025-08-07 |
| 43 | GPT-5.1-Codex-Max | 9 | unverifiedT1 | 2025-11-19 |
| 44 | DeepSeek-V3.2 | 8 | unverifiedT1 | 2025-12-01 |
| 45 | Inkling-Small | 6 | unverifiedT1 | 2026-07-30 |
| 46 | GLM-4.7 | 6 | unverifiedT1 | 2025-12-22 |
| 47 | GPT-5.4 Nano | 5 | unverifiedT1 | 2026-03-17 |
| 48 | MiniMax-M2.5 | 4 | unverifiedT1 | 2026-02-12 |
| 49 | MiniMax-M2.7 | 3 | unverifiedT1 | 2026-03-18 |
| 50 | Nemotron 3 Ultra | 2 | unverifiedT1 | 2026-06-04 |
| 51 | Inkling | 0 | unverifiedT1 | 2026-07-15 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
This benchmark evaluates 51 models, and the score difference between the best and worst result is 99.0 points. A spread of this width indicates a genuine capability gap separating the models, rather than a narrow margin where rank differences could be accidental.
3 cited facts
The benchmark's Benchmark Integrity Index is 99, its grade is A, and it ranks 12th out of 61 benchmarks. Saturation is the weakest integrity component, meaning the ceiling is crowded for top scores. This index indicates the benchmark's health as a discriminator of frontier models under a disclosed harness, not a statement about any model's capability.
6 cited facts
The benchmark is not saturated, with a top score of 99.0 and only 1.0 points of remaining headroom to the ceiling. Only one model clusters near the top, so there is still genuine room to separate the field at the top; small differences are not yet noise-bound, though the remaining separation is within a couple of points.
4 cited facts
This benchmark is evaluated with a consistent harness, so scores are directly comparable when the same harness is used. Because the test set privacy is unknown, a public test set would warrant heightened scrutiny of high scores for possible contamination, while a held-out set would reduce that concern. Contamination history is unknown, so no qualitative contamination risk can be inferred; scores should be interpreted as task performance under the disclosed harness, not as deployed capability.
4 cited facts
This benchmark's intended measurement is not documented, so its precise scope cannot be described from the available facts. Its construction and key assumptions are also undocumented, leaving the design rationale unavailable. The sharpest caveat is that because limitations are not documented, any claim about what this benchmark cannot tell you would be unsupported.
4 cited facts
ProofBench: ProofBench as reported in Epoch AI's Capabilities Index CSV.
Claude Opus 5 leads ProofBench at 99. The full leaderboard above lists every recorded measurement, not just the headline number.
51 models have recorded scores on ProofBench, spanning a score spread of 99.
tensor.news grades ProofBench A for integrity (score 99/100), ranking #12 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — ProofBench still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is unknown, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Scores are reported under a consistent harness, so comparisons on ProofBench are reasonably apples-to-apples.
Every ProofBench measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites epoch.ai/data/eci_benchmarks.csv.