ARC-AGI-2
Integrity rank #1 of 61 · 71 models scored · top score 92.5 · GPT-5.6 Sol
Fluid intelligence / novel skill via abstract grid puzzles
Strongest on discrimination, weakest on saturation headroom.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | GPT-5.6 Sol | 92.5 | unverifiedT2 | 2026-07-09 |
| 2 | Claude Opus 5 | 90.42 | unverifiedT2 | 2026-07-24 |
| 3 | Claude Fable 5 | 89.17 | unverifiedT2 | 2026-06-09 |
| 4 | GPT-5.5 | 85 | unverifiedT2 | 2026-04-23 |
| 5 | GPT-5.5 Pro | 84.58 | unverifiedT2 | 2026-04-23 |
| 6 | GPT-5.6 Terra | 83.9 | unverifiedT2 | 2026-07-09 |
| 7 | GPT-5.4 Pro | 83.33 | unverifiedT2 | 2026-03-05 |
| 8 | Gemini 3.1 Pro | 77.1 | unverifiedT2 | 2026-02-19 |
| 9 | Claude Opus 4.7 | 75.83 | unverifiedT2 | 2026-04-16 |
| 10 | GPT-5.4 | 73.95 | unverifiedT2 | 2026-03-05 |
| 11 | Gemini 3.5 Flash | 72.08 | unverifiedT2 | 2026-05-19 |
| 12 | Claude Opus 4.8 | 72.08 | unverifiedT2 | 2026-05-28 |
| 13 | Claude Opus 4.6 | 69.17 | unverifiedT2 | 2026-02-05 |
| 14 | Grok 4.6 | 67.08 | unverifiedT2 | 2026-08-12 |
| 15 | Grok 4.20 | 65.14 | unverifiedT2 | 2026-02-17 |
| 16 | DeepSeek V4 Flash 0731 | 61.39 | unverifiedT2 | 2026-07-31 |
| 17 | Gemini 3.6 Flash | 60.42 | unverifiedT2 | 2026-07-21 |
| 18 | Kimi K3 | 60.42 | unverifiedT2 | 2026-07-16 |
| 19 | Claude Sonnet 4.6 | 60.42 | unverifiedT2 | 2026-02-17 |
| 20 | GPT-5.6 Luna | 59.54 | unverifiedT2 | 2026-07-09 |
| 21 | GPT-5.2 Pro | 54.16 | unverifiedT2 | 2025-12-11 |
| 22 | GPT-5.2 | 52.91 | unverifiedT2 | 2025-12-11 |
| 23 | Grok 4.5 | 52.64 | unverifiedT2 | 2026-07-08 |
| 24 | Inkling-Small | 40.14 | unverifiedT2 | 2026-07-30 |
| 25 | Claude Opus 4.5 | 37.64 | unverifiedT2 | 2025-11-24 |
| 26 | Inkling | 36.53 | unverifiedT2 | 2026-07-15 |
| 27 | Gemini 3 Flash | 33.61 | unverifiedT2 | 2025-12-17 |
| 28 | Gemini 3 Pro | 31.11 | unverifiedT2 | 2025-11-18 |
| 29 | GLM-5.2 | 22.78 | unverifiedT2 | 2026-06-16 |
| 30 | GPT-5.4 Mini | 18.9 | unverifiedT2 | 2026-03-17 |
| 31 | GPT-5 Pro | 18.33 | unverifiedT2 | 2025-10-07 |
| 32 | GPT-5.1 | 17.64 | unverifiedT2 | 2025-11-13 |
| 33 | Grok 4 | 15.97 | unverifiedT2 | 2025-07-09 |
| 34 | Claude Sonnet 4.5 | 13.61 | unverifiedT2 | 2025-09-29 |
| 35 | Kimi K2.5 | 11.81 | unverifiedT2 | 2026-02-02 |
| 36 | Gemini 3.5 Flash-Lite | 10.28 | unverifiedT2 | 2026-07-21 |
| 37 | GPT-5 | 9.86 | unverifiedT2 | 2025-08-07 |
| 38 | Claude Opus 4 | 8.61 | unverifiedT2 | 2025-05-22 |
| 39 | o3 | 6.53 | unverifiedT2 | 2024-12-20 |
| 40 | o4-mini | 6.11 | unverifiedT2 | 2025-04-16 |
| 41 | Claude Sonnet 4 | 5.93 | unverifiedT2 | 2025-05-22 |
| 42 | GPT-5.4 Nano | 5.69 | unverifiedT2 | 2026-03-17 |
| 43 | Grok 4 Fast | 5.28 | unverifiedT2 | 2025-09-19 |
| 44 | GLM-5 | 4.86 | unverifiedT2 | 2026-02-11 |
| 45 | MiniMax-M2.5 | 4.86 | unverifiedT2 | 2026-02-12 |
| 46 | Gemini 2.5 Pro (Jun 2025) | 4.86 | unverifiedT2 | 2025-06-05 |
| 47 | o3-pro | 4.86 | unverifiedT2 | 2025-06-10 |
| 48 | GPT-5 mini | 4.44 | unverifiedT2 | 2025-08-07 |
| 49 | DeepSeek-V3.2 | 4.03 | unverifiedT2 | 2025-12-01 |
| 50 | Claude Haiku 4.5 | 4.03 | unverifiedT2 | 2025-10-15 |
| 51 | o3-mini | 2.99 | unverifiedT2 | 2025-01-31 |
| 52 | GPT-5 nano | 2.61 | unverifiedT2 | 2025-08-07 |
| 53 | Gemini 2.5 Flash (May 2025) | 2.54 | unverifiedT2 | 2025-05-20 |
| 54 | DeepSeek-R1 | 1.3 | unverifiedT2 | 2025-01-20 |
| 55 | Gemini 2.0 Flash (Feb 2025) | 1.3 | unverifiedT2 | 2024-12-11 |
| 56 | Qwen3-235B-A22B-Instruct (Jul 2025) | 1.25 | unverifiedT2 | 2025-07-25 |
| 57 | DeepSeek-R1 (May 2025) | 1.12 | unverifiedT2 | 2025-05-28 |
| 58 | Claude 3.7 Sonnet | 0.9 | unverifiedT2 | 2025-02-24 |
| 59 | o1-mini | 0.83 | unverifiedT2 | 2024-09-12 |
| 60 | GPT-4.5 | 0.8 | unverifiedT2 | 2025-02-27 |
| 61 | Gemini 1.5 Pro (Sept 2024) | 0.8 | unverifiedT2 | 2024-09-24 |
| 62 | Grok-3 mini | 0.42 | unverifiedT2 | 2025-02-19 |
| 63 | GPT-4.1 | 0.42 | unverifiedT2 | 2025-04-14 |
| 64 | Grok 3 | 0 | unverifiedT2 | 2025-02-17 |
| 65 | Llama 4 Maverick | 0 | unverifiedT2 | 2025-04-05 |
| 66 | Llama 4 Scout | 0 | unverifiedT2 | 2025-04-05 |
| 67 | Magistral Small 1.0 | 0 | unverifiedT2 | 2025-06-10 |
| 68 | GPT-4.1 mini | 0 | unverifiedT2 | 2025-04-14 |
| 69 | GPT-4.1 nano | 0 | unverifiedT2 | 2025-04-14 |
| 70 | GPT-4o (Nov 2024) | 0 | unverifiedT2 | 2024-05-13 |
| 71 | GPT-4o mini | 0 | unverifiedT2 | 2024-07-18 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
This benchmark scores 71 models, with a 92.5-point spread between the best and worst results; that wide margin indicates the observed score differences likely reflect genuine capability separation rather than noise.
2 cited facts
The benchmark has an integrity score of 100 with an integrity grade of A, ranking 1st among 61 benchmarks in benchmark universe size, reflecting strong health as a discriminator of frontier models under a disclosed harness. Its weakest component is saturation, which means the ceiling is crowded and top-score differences become harder to interpret.
5 cited facts
The benchmark is not saturated, with a top score of 92.5 and remaining headroom of 7.5 points to the ceiling. Only one model sits near the top, all within a couple of points of the leading score, so the field at the top is not yet crowded. Because the benchmark is not saturated, there is still genuine room to separate the top models, so current differences near the ceiling are meaningful rather than noise.
5 cited facts
Task performance is measured under a consistent harness, so scores are directly comparable within this evaluation framework. The test set is held out rather than public, reducing contamination risk relative to open test sets; indeed, the private eval has not been released and is low contamination by design. This harness also shows extreme test-time-compute sensitivity with a pass-based metric, so scores should be read as task performance under this disclosed setup rather than as deployed capability.
4 cited facts
This benchmark targets fluid intelligence and novel skill acquisition through abstract grid puzzles, using tasks that are not drawn from any existing corpus. Its test items are hand-designed grid tasks calibrated against human performance, split into a public train/eval set and a secret eval set, on the assumption that human-solvable but model-hard problems isolate general reasoning. The sharpest caveat is that the reported score reflects search and compute budget, and the narrow grid domain means the benchmark can mislead if interpreted as a broad measure of reasoning rather than performance in this controlled setting.
5 cited facts
ARC-AGI-2: ARC-AGI-2 — the second-generation Abstraction and Reasoning Corpus of novel visual-grid puzzles measuring fluid, few-shot abstraction that resists brute-force memorization; scored as the percent of tasks solved. The evaluation set is held out (private) to keep it contamination-resistant.
GPT-5.6 Sol leads ARC-AGI-2 at 92.5. The full leaderboard above lists every recorded measurement, not just the headline number.
71 models have recorded scores on ARC-AGI-2, spanning a score spread of 92.5.
tensor.news grades ARC-AGI-2 A for integrity (score 100/100), ranking #1 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — ARC-AGI-2 still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is held-out, which lowers — but does not eliminate — the risk that training data leaked into the questions.
Scores are reported under a consistent harness, so comparisons on ARC-AGI-2 are reasonably apples-to-apples.
Every ARC-AGI-2 measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites arxiv.org/abs/2505.11831.