Benchmark explainer
What is ARC-AGI-2?
Fluid intelligence / novel skill via abstract grid puzzles
ARC-AGI-2 — the second-generation Abstraction and Reasoning Corpus of novel visual-grid puzzles measuring fluid, few-shot abstraction that resists brute-force memorization; scored as the percent of tasks solved. The evaluation set is held out (private) to keep it contamination-resistant.
How it's scored
- Metric
- pass@2 accuracy
- Score ceiling
- 100
- Construction
- Hand-designed grid tasks, human-calibrated; public train/eval + secret eval set
- Human baseline
- Human panel 100%; avg individual ~60%
How trustworthy
Discrimination
100/100Does it still separate models?
Saturation headroom
98.7/100How far from ceiling / clustered at the top?
Contamination resistance
100/100Public vs held-out; training-leak risk.
Harness comparability
100/100Is it apples-to-apples, or mixed / vendor-optimized?
Freshness
100/100How old is the benchmark?
Contamination history: Private eval not released; low contamination by design
Limitations: Narrow grid domain; program-search brute-forces; score reflects search+compute budget
Who leads ARC-AGI-2
| Model | Score | Evidence |
|---|---|---|
| GPT-5.6 Sol | 92.5 | unverified |
| Claude Opus 5 | 90.42 | unverified |
| Claude Fable 5 | 89.17 | unverified |
| GPT-5.5 | 85 | unverified |
| GPT-5.5 Pro | 84.58 | unverified |
Frequently asked questions
ARC-AGI-2: ARC-AGI-2 — the second-generation Abstraction and Reasoning Corpus of novel visual-grid puzzles measuring fluid, few-shot abstraction that resists brute-force memorization; scored as the percent of tasks solved. The evaluation set is held out (private) to keep it contamination-resistant.
GPT-5.6 Sol leads ARC-AGI-2 at 92.5. The full leaderboard above lists every recorded measurement, not just the headline number.
71 models have recorded scores on ARC-AGI-2, spanning a score spread of 92.5.
tensor.news grades ARC-AGI-2 A for integrity (score 100/100), ranking #1 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — ARC-AGI-2 still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is held-out, which lowers — but does not eliminate — the risk that training data leaked into the questions.
Scores are reported under a consistent harness, so comparisons on ARC-AGI-2 are reasonably apples-to-apples.
Every ARC-AGI-2 measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites arxiv.org/abs/2505.11831.