Benchmark explainer
What is ARC-AGI?
unknown
ARC-AGI as reported in Epoch AI's Capabilities Index CSV.
How it's scored
- Metric
- normalized performance in Epoch CSV, rendered as accuracy (%)
- Score ceiling
- 100
- Construction
- unknown
- Human baseline
- unknown
How trustworthy
Discrimination
100/100Does it still separate models?
Saturation headroom
91.8/100How far from ceiling / clustered at the top?
Contamination resistance
0/100Public vs held-out; training-leak risk.
Harness comparability
100/100Is it apples-to-apples, or mixed / vendor-optimized?
Freshness
0/100How old is the benchmark?
Contamination history: unknown
Limitations: unknown
Who leads ARC-AGI
| Model | Score | Evidence |
|---|---|---|
| Claude Fable 5 | 98.5 | unverified |
| Gemini 3.1 Pro | 98 | unverified |
| Claude Opus 5 | 97.5 | unverified |
| GPT-5.6 Sol | 97.5 | unverified |
| GPT-5.5 Pro | 96.5 | unverified |
Frequently asked questions
ARC-AGI: ARC-AGI as reported in Epoch AI's Capabilities Index CSV.
Claude Fable 5 leads ARC-AGI at 98.5. The full leaderboard above lists every recorded measurement, not just the headline number.
72 models have recorded scores on ARC-AGI, spanning a score spread of 98.5.
tensor.news grades ARC-AGI B for integrity (score 73/100), ranking #47 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
Yes — top scores are clustering near the ceiling (saturation ratio 0.08), so ARC-AGI no longer separates leading models well. Treat small gaps at the top with caution.
Its test set is unknown, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Scores are reported under a consistent harness, so comparisons on ARC-AGI are reasonably apples-to-apples.
Every ARC-AGI measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites epoch.ai/data/eci_benchmarks.csv.