Anthropic·released 2026-06-091 source
Claude Fable 5 benchmark scores: 21 benchmarks tracked, leading 11. 43% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 11 of 21 benchmarks·best result 100 on OTIS Mock AIME 2024-2025 (reproduced)·9 of 21 independently reproduced·$10/$50 per M tokens
Head-to-headClaude Fable 5 vs Claude Opus 5Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
Leads on: APEX-Agents, ARC-AGI, FrontierCode, FrontierMath-Tier-4-v2-Private, MirrorCode, OTIS Mock AIME 2024-2025, PostTrainBench, Remote Labor Index, SimpleBench, Surface Evolver Bench, WeirdML
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 100accuracy (%) | reproduced· optimizedT1 | 2026-06-09Epoch AI |
| ARC-AGI | 98.5accuracy (%) | unverified· optimizedT1 | 2026-06-09https://arcprize.org/leaderboard |
| Surface Evolver Bench | 95accuracy (%) | unverifiedT1 | 2026-06-09https://yhenon.github.io/surface-evolver-llm-eval/ |
| ProofBench | 95accuracy (%) | unverifiedT1 | 2026-06-09https://www.vals.ai/benchmarks/proof_bench |
| WeirdML | 91.94accuracy (%) | unverifiedT1 | 2026-06-09https://htihle.github.io/weirdml.html |
| ARC-AGI-2 | 89.17accuracy (%) | unverifiedT2 | 2026-06-09 |
| FrontierMath-Tier-4-v2-Private | 87.8accuracy (%) | reproducedT1 | 2026-06-09Epoch AI |
| FrontierMath-Tiers-1-3-v2-Private | 87.02accuracy (%) | reproducedT1 | 2026-06-09Epoch AI |
| GPQA diamond | 81.14accuracy (%) | reproduced· optimizedT1 | 2026-06-09Epoch AI |
| SimpleBench | 78.28accuracy (%) | unverifiedT1 | 2026-06-09SimpleBench Leaderboard |
| DeepSWE | 69.91accuracy (%) | unverifiedT1 | 2026-06-09https://deepswe.datacurve.ai/ |
| SimpleQA Verified | 68.3accuracy (%) | reproducedT1 | 2026-06-09Epoch AI |
| MirrorCode | 63.89accuracy (%) | reproducedT1 | 2026-06-09Epoch AI |
| FrontierCode | 53.5accuracy (%) | unverifiedT1 | 2026-06-09https://cognition.com/frontiercode |
| Mystery Game Puzzles | 47.12accuracy (%) | reproducedT1 | 2026-06-09Epoch AI |
| APEX-Agents | 45accuracy (%) | unverifiedT2 | 2026-06-09 |
| PostTrainBench | 41.79accuracy (%) | unverifiedT2 | 2026-06-09 |
| EBR-bench | 39.52accuracy (%) | reproducedT1 | 2026-06-09Epoch AI |
| Chess Puzzles | 37.92accuracy (%) | reproducedT1 | 2026-06-09Epoch AI |
| CritPt | 28.57accuracy (%) | unverifiedT2 | 2026-06-09 |
| Remote Labor Index | 16.1accuracy (%) | unverifiedT2 | 2026-06-09 |
Claims drawn from cited facts, not live model generation.
This model is scored on 21 benchmarks and is currently at the top of 11 of them; because the record has been independently reproduced, outside evaluation backs these scores up, so the results can be trusted.
3 cited facts
On average this model trails the leader by 3.82 points, a moderate gap that reads as competitive but not front-of-pack. It holds the top score on 11 benchmarks, indicating genuine category leadership on those tasks despite the overall gap.
2 cited facts
6 facts cross-checked across data sources: 3 corroborated, 2 single-source, 1 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: curated_capability_claims, epoch_benchmarks, litellm_prices, modelsdev_models +1 more
Claude Fable 5 is an AI model developed by Anthropic, released 2026-06-09. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Claude Fable 5 has recorded scores on 21 benchmarks, each shown with its evidence status.
Claude Fable 5 has recorded scores on 21 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleBench, SimpleQA Verified, and 15 more. The full table above shows each score with its evidence status.
9 of 21 recorded scores (43%) are independently reproduced rather than self-reported by the lab.
Claude Fable 5 has 21 tracked claims: 9 independently reproduced, 12 unverified.
Listed API pricing: $10 per million input tokens, $50 per million output tokens. See the pricing block for the full breakdown.