Skip to content

Benchmark explainer

What is ARC-AGI?

unknown

ARC-AGI as reported in Epoch AI's Capabilities Index CSV.

BINTEGRITY 73 / 100tensor.news
consistent harnesspublic test setsaturated

How it's scored

Metric
normalized performance in Epoch CSV, rendered as accuracy (%)
Score ceiling
100
Construction
unknown
Human baseline
unknown

How trustworthy

Discrimination

100/100

Does it still separate models?

Saturation headroom

91.8/100

How far from ceiling / clustered at the top?

Contamination resistance

0/100

Public vs held-out; training-leak risk.

Harness comparability

100/100

Is it apples-to-apples, or mixed / vendor-optimized?

Freshness

0/100

How old is the benchmark?

Contamination history: unknown

Limitations: unknown

Who leads ARC-AGI

ModelScoreEvidence
Claude Fable 598.5unverified
Gemini 3.1 Pro98unverified
Claude Opus 597.5unverified
GPT-5.6 Sol97.5unverified
GPT-5.5 Pro96.5unverified

Frequently asked questions

ARC-AGI: ARC-AGI as reported in Epoch AI's Capabilities Index CSV.

Claude Fable 5 leads ARC-AGI at 98.5. The full leaderboard above lists every recorded measurement, not just the headline number.

72 models have recorded scores on ARC-AGI, spanning a score spread of 98.5.

tensor.news grades ARC-AGI B for integrity (score 73/100), ranking #47 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.

Yes — top scores are clustering near the ceiling (saturation ratio 0.08), so ARC-AGI no longer separates leading models well. Treat small gaps at the top with caution.

Its test set is unknown, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.

Scores are reported under a consistent harness, so comparisons on ARC-AGI are reasonably apples-to-apples.

Every ARC-AGI measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites epoch.ai/data/eci_benchmarks.csv.

Follow the record