ARC-AGI
Integrity rank #47 of 61 · 72 models scored · top score 98.5 · Claude Fable 5
unknown
Strongest on discrimination, weakest on contamination resistance. Caveat: the public test set does not rule out training-data contamination.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | Claude Fable 5 | 98.5 | unverified· optimizedT1 | 2026-06-09 |
| 2 | Gemini 3.1 Pro | 98 | unverified· optimizedT1 | 2026-02-19 |
| 3 | Claude Opus 5 | 97.5 | unverified· optimizedT1 | 2026-07-24 |
| 4 | GPT-5.6 Sol | 97.5 | unverified· optimizedT1 | 2026-07-09 |
| 5 | GPT-5.5 Pro | 96.5 | unverified· optimizedT1 | 2026-04-23 |
| 6 | GPT-5.6 Terra | 96.5 | unverified· optimizedT1 | 2026-07-09 |
| 7 | GPT-5.5 | 95 | unverified· optimizedT1 | 2026-04-23 |
| 8 | Kimi K3 | 94.5 | unverified· optimizedT1 | 2026-07-16 |
| 9 | GPT-5.4 Pro | 94.5 | unverified· optimizedT1 | 2026-03-05 |
| 10 | Claude Opus 4.6 | 94 | unverified· optimizedT1 | 2026-02-05 |
| 11 | GPT-5.4 | 93.67 | unverified· optimizedT1 | 2026-03-05 |
| 12 | Claude Opus 4.7 | 93.5 | unverified· optimizedT1 | 2026-04-16 |
| 13 | Gemini 3.5 Flash | 92.5 | unverified· optimizedT1 | 2026-05-19 |
| 14 | Claude Opus 4.8 | 92.5 | unverified· optimizedT1 | 2026-05-28 |
| 15 | Gemini 3.6 Flash | 91.17 | unverified· optimizedT1 | 2026-07-21 |
| 16 | GPT-5.2 Pro | 90.5 | unverified· optimizedT1 | 2025-12-11 |
| 17 | Grok 4.20 | 89.5 | unverified· optimizedT1 | 2026-02-17 |
| 18 | DeepSeek V4 Flash 0731 | 89 | unverified· optimizedT1 | 2026-07-31 |
| 19 | GPT-5.6 Luna | 88 | unverified· optimizedT1 | 2026-07-09 |
| 20 | Grok 4.6 | 87.5 | unverified· optimizedT1 | 2026-08-12 |
| 21 | Grok 4.5 | 87.17 | unverified· optimizedT1 | 2026-07-08 |
| 22 | Claude Sonnet 4.6 | 86.5 | unverified· optimizedT2 | 2026-02-17 |
| 23 | GPT-5.2 | 86.2 | unverified· optimizedT1 | 2025-12-11 |
| 24 | Gemini 3 Flash | 84.67 | unverified· optimizedT1 | 2025-12-17 |
| 25 | Inkling-Small | 84 | unverified· optimizedT1 | 2026-07-30 |
| 26 | Claude Opus 4.5 | 80 | unverified· optimizedT1 | 2025-11-24 |
| 27 | Inkling | 79.5 | unverified· optimizedT1 | 2026-07-15 |
| 28 | GLM-5.2 | 77 | unverified· optimizedT1 | 2026-06-16 |
| 29 | Gemini 3 Pro | 75 | unverified· optimizedT1 | 2025-11-18 |
| 30 | GPT-5.1 | 72.83 | unverified· optimizedT1 | 2025-11-13 |
| 31 | GPT-5 Pro | 70.2 | unverified· optimizedT1 | 2025-10-07 |
| 32 | Grok 4 | 66.67 | unverified· optimizedT2 | 2025-07-09 |
| 33 | GPT-5 | 65.7 | unverified· optimizedT1 | 2025-08-07 |
| 34 | Kimi K2.5 | 65.33 | unverified· optimizedT2 | 2026-02-02 |
| 35 | MiniMax-M2.5 | 63.67 | unverified· optimizedT2 | 2026-02-12 |
| 36 | Claude Sonnet 4.5 | 63.67 | unverified· optimizedT2 | 2025-09-29 |
| 37 | GPT-5.4 Mini | 63.67 | unverified· optimizedT1 | 2026-03-17 |
| 38 | o3 | 60.8 | unverified· optimizedT1 | 2024-12-20 |
| 39 | o3-pro | 59.3 | unverified· optimizedT1 | 2025-06-10 |
| 40 | o4-mini | 58.7 | unverified· optimizedT1 | 2025-04-16 |
| 41 | DeepSeek-V3.2 | 57 | unverified· optimizedT2 | 2025-12-01 |
| 42 | GPT-5 mini | 54.33 | unverified· optimizedT2 | 2025-08-07 |
| 43 | Gemini 3.5 Flash-Lite | 53.5 | unverified· optimizedT1 | 2026-07-21 |
| 44 | GPT-5.4 Nano | 51.5 | unverified· optimizedT1 | 2026-03-17 |
| 45 | Grok 4 Fast | 48.5 | unverified· optimizedT2 | 2025-09-19 |
| 46 | Claude Haiku 4.5 | 47.67 | unverified· optimizedT2 | 2025-10-15 |
| 47 | GLM-5 | 44.67 | unverified· optimizedT2 | 2026-02-11 |
| 48 | Gemini 2.5 Pro (Jun 2025) | 41 | unverified· optimizedT1 | 2025-06-05 |
| 49 | Claude Sonnet 4 | 40 | unverified· optimizedT1 | 2025-05-22 |
| 50 | Claude Opus 4 | 35.7 | unverified· optimizedT1 | 2025-05-22 |
| 51 | o3-mini | 34.5 | unverified· optimizedT1 | 2025-01-31 |
| 52 | Gemini 2.5 Flash (May 2025) | 33.3 | unverified· optimizedT1 | 2025-05-20 |
| 53 | Gemini 2.5 Pro (Mar 2025) | 33 | unverified· optimizedT1 | 2025-03-25 |
| 54 | Gemini 2.5 Flash (Apr 2025) | 32.3 | unverified· optimizedT1 | 2025-04-17 |
| 55 | o1 | 30.7 | unverified· optimizedT1 | 2024-12-05 |
| 56 | Claude 3.7 Sonnet | 28.6 | unverified· optimizedT1 | 2025-02-24 |
| 57 | DeepSeek-R1 (May 2025) | 21.2 | unverified· optimizedT1 | 2025-05-28 |
| 58 | GPT-5 nano | 20.71 | unverified· optimizedT2 | 2025-08-07 |
| 59 | o1-preview | 18 | unverified· optimizedT1 | 2024-09-12 |
| 60 | Grok-3 mini | 16.5 | unverified· optimizedT1 | 2025-02-19 |
| 61 | DeepSeek-R1 | 15.8 | unverified· optimizedT1 | 2025-01-20 |
| 62 | o1-mini | 14 | unverified· optimizedT1 | 2024-09-12 |
| 63 | Qwen3-235B-A22B-Instruct (Jul 2025) | 11 | unverified· optimizedT2 | 2025-07-25 |
| 64 | GPT-4.5 | 10.3 | unverified· optimizedT1 | 2025-02-27 |
| 65 | Grok 3 | 5.5 | unverified· optimizedT1 | 2025-02-17 |
| 66 | GPT-4.1 | 5.5 | unverified· optimizedT1 | 2025-04-14 |
| 67 | Magistral Small 1.0 | 5 | unverified· optimizedT2 | 2025-06-10 |
| 68 | GPT-4o (Nov 2024) | 4.5 | unverified· optimizedT1 | 2024-05-13 |
| 69 | Llama 4 Maverick | 4.4 | unverified· optimizedT1 | 2025-04-05 |
| 70 | GPT-4.1 mini | 3.5 | unverified· optimizedT1 | 2025-04-14 |
| 71 | Llama 4 Scout | 0.5 | unverified· optimizedT1 | 2025-04-05 |
| 72 | GPT-4.1 nano | 0 | unverified· optimizedT1 | 2025-04-14 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
This benchmark scores 72 models, and the best result exceeds the worst by 98.5 points. This wide spread supports real separation between models, so large rank differences are likely to reflect a genuine capability gap.
3 cited facts
This benchmark receives an integrity score of 73 (grade B), ranking 47th out of 61 benchmarks in the universe. Its weakest integrity component is contamination, which narrows how much readers should trust top scores; the score reflects the benchmark's health as a frontier-model discriminator under a disclosed harness, not any model's capability.
5 cited facts
With a top score of 98.5, only 1.5 points of headroom, and six models within a couple of points of the top, this benchmark is saturated. For the reader, saturation means today's ranking among the top models is close to noise, so do not read small differences there as real capability gaps.
5 cited facts
Scores in this section were produced under a consistent harness, so task performance is directly comparable for head-to-head comparisons within this benchmark. The test set is neither explicitly public nor held out, so high scores should be treated cautiously because the possibility of contamination cannot be assessed. Any contamination history is undocumented, meaning the absence of a reported contamination window should not be read as evidence of cleanliness.
3 cited facts
This benchmark's intended measurement is not documented, as what it measures is unknown. Its construction and key assumptions are likewise undocumented, so the design rationale cannot be reconstructed from the available facts. The sharpest caveat is that, with limitations and any human baseline unspecified, this benchmark offers no documented account of what it fails to capture or where it is likely to mislead.
5 cited facts
ARC-AGI: ARC-AGI as reported in Epoch AI's Capabilities Index CSV.
Claude Fable 5 leads ARC-AGI at 98.5. The full leaderboard above lists every recorded measurement, not just the headline number.
72 models have recorded scores on ARC-AGI, spanning a score spread of 98.5.
tensor.news grades ARC-AGI B for integrity (score 73/100), ranking #47 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
Yes — top scores are clustering near the ceiling (saturation ratio 0.08), so ARC-AGI no longer separates leading models well. Treat small gaps at the top with caution.
Its test set is unknown, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Scores are reported under a consistent harness, so comparisons on ARC-AGI are reasonably apples-to-apples.
Every ARC-AGI measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites epoch.ai/data/eci_benchmarks.csv.