Anthropic·released 2026-09-011 source
Claude Fable 5.1 benchmark scores: 15 benchmarks tracked, leading 6. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 6 of 15 benchmarks·best result 100 on OTIS Mock AIME 2024-2025 (unverified)·0 of 15 independently reproduced
Head-to-headClaude Fable 5.1 vs GPT-6 AstraLeads on: APEX-Agents, HLE, MirrorCode, OTIS Mock AIME 2024-2025, ProofBench, WeirdML
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 100accuracy (%) | unverifiedT2 | 2026-09-01 |
| ProofBench | 100accuracy (%) | unverifiedT1 | 2026-09-01https://www.vals.ai/benchmarks/proof_bench |
| ARC-AGI | 97.5accuracy (%) | unverifiedT1 | 2026-09-01https://arcprize.org/leaderboard |
| WeirdML | 92.9accuracy (%) | unverifiedT1 | 2026-09-01https://htihle.github.io/weirdml.html |
| FrontierMath-Tiers-1-3-v2-Private | 90.18accuracy (%) | unverifiedT2 | 2026-09-01 |
| ARC-AGI-2 | 90accuracy (%) | unverifiedT2 | 2026-09-01 |
| FrontierMath-Tier-4-v2-Private | 87.8accuracy (%) | unverifiedT2 | 2026-09-01 |
| MirrorCode | 73.3accuracy (%) | unverifiedT1 | 2026-09-01Epoch evaluations (staging) |
| SimpleQA Verified | 70.8accuracy (%) | unverifiedT2 | 2026-09-01 |
| EBR-bench | 57.14accuracy (%) | unverifiedT2 | 2026-09-01 |
| Mystery Game Puzzles | 53.73accuracy (%) | unverifiedT2 | 2026-09-01 |
| FrontierCode | 50.91accuracy (%) | unverifiedT1 | 2026-09-01https://cognition.com/frontiercode |
| APEX-Agents | 47.4accuracy (%) | unverifiedT2 | 2026-09-01 |
| Chess Puzzles | 44.23accuracy (%) | unverifiedT2 | 2026-09-01 |
| HLE | 43.8accuracy (%) | unverifiedT2 | 2026-09-01 |
Claims drawn from cited facts, not live model generation.
This model was developed by Anthropic. It was released in 2026, making it a relatively recent arrival in that developer's lineup.
3 cited facts
It is scored on 15 tracked benchmarks, and of those it currently tops 6. Its record is unverified: neither confirmed nor disputed by any outside evaluation, so the numbers should be read as reported figures rather than settled results.
3 cited facts
This model trails the state-of-the-art leader by an average of 6.71 points across its benchmarks, a moderate gap that reads as competitive but not front-of-pack. It holds the top score on 6 benchmarks, indicating some category leadership.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Claude Fable 5.1 is an AI model developed by Anthropic, released 2026-09-01. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Claude Fable 5.1 has recorded scores on 15 benchmarks, each shown with its evidence status.
Claude Fable 5.1 has recorded scores on 15 benchmarks — OTIS Mock AIME 2024-2025, WeirdML, SimpleQA Verified, Chess Puzzles, ARC-AGI, ARC-AGI-2, and 9 more. The full table above shows each score with its evidence status.
0 of 15 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Claude Fable 5.1 has 15 tracked claims: 15 unverified.