Skip to content

Benchmark explainer

What is ARC-AGI-2?

Fluid intelligence / novel skill via abstract grid puzzles

ARC-AGI-2 — the second-generation Abstraction and Reasoning Corpus of novel visual-grid puzzles measuring fluid, few-shot abstraction that resists brute-force memorization; scored as the percent of tasks solved. The evaluation set is held out (private) to keep it contamination-resistant.

AINTEGRITY 100 / 100tensor.news
consistent harnessheld-out test set

How it's scored

Metric
pass@2 accuracy
Score ceiling
100
Construction
Hand-designed grid tasks, human-calibrated; public train/eval + secret eval set
Human baseline
Human panel 100%; avg individual ~60%

How trustworthy

Discrimination

100/100

Does it still separate models?

Saturation headroom

98.7/100

How far from ceiling / clustered at the top?

Contamination resistance

100/100

Public vs held-out; training-leak risk.

Harness comparability

100/100

Is it apples-to-apples, or mixed / vendor-optimized?

Freshness

100/100

How old is the benchmark?

Contamination history: Private eval not released; low contamination by design

Limitations: Narrow grid domain; program-search brute-forces; score reflects search+compute budget

Who leads ARC-AGI-2

ModelScoreEvidence
GPT-5.6 Sol92.5unverified
Claude Opus 590.42unverified
Claude Fable 589.17unverified
GPT-5.585unverified
GPT-5.5 Pro84.58unverified

Frequently asked questions

ARC-AGI-2: ARC-AGI-2 — the second-generation Abstraction and Reasoning Corpus of novel visual-grid puzzles measuring fluid, few-shot abstraction that resists brute-force memorization; scored as the percent of tasks solved. The evaluation set is held out (private) to keep it contamination-resistant.

GPT-5.6 Sol leads ARC-AGI-2 at 92.5. The full leaderboard above lists every recorded measurement, not just the headline number.

71 models have recorded scores on ARC-AGI-2, spanning a score spread of 92.5.

tensor.news grades ARC-AGI-2 A for integrity (score 100/100), ranking #1 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.

No — ARC-AGI-2 still has headroom and continues to discriminate between models rather than bunching them at the ceiling.

Its test set is held-out, which lowers — but does not eliminate — the risk that training data leaked into the questions.

Scores are reported under a consistent harness, so comparisons on ARC-AGI-2 are reasonably apples-to-apples.

Every ARC-AGI-2 measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites arxiv.org/abs/2505.11831.

Follow the record