Anthropic·released 2026-04-161 source
Claude Opus 4.7 benchmark scores: 24 benchmarks tracked, leading 1. 42% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 1 of 24 benchmarks·best result 97.8 on OTIS Mock AIME 2024-2025 (reproduced)·10 of 24 independently reproduced·$5/$25 per M tokens
Head-to-headClaude Opus 4.7 vs GPT-5.5Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
Leads on: SWE-Bench verified
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 97.8accuracy (%) | reproduced· optimizedT1 | 2026-04-16Epoch AI |
| ARC-AGI | 93.5accuracy (%) | unverified· optimizedT1 | 2026-04-16https://arcprize.org/leaderboard |
| GPQA diamond | 86.87accuracy (%) | reproduced· optimizedT1 | 2026-04-16Epoch AI |
| SWE-Bench verified | 83.47accuracy (%) | reproduced· optimizedT1 | 2026-04-16Epoch AI |
| Terminal Bench | 80.2accuracy (%) | unverified· optimizedT1 | 2026-04-16https://www.tbench.ai/leaderboard/terminal-bench/2.0 |
| WeirdML | 76.4accuracy (%) | unverifiedT1 | 2026-04-16https://htihle.github.io/weirdml.html |
| ARC-AGI-2 | 75.83accuracy (%) | unverifiedT2 | 2026-04-16 |
| FrontierMath-Tiers-1-3-v2-Private | 70.18accuracy (%) | reproducedT1 | 2026-04-16Epoch AI |
| SimpleBench | 55.48accuracy (%) | unverifiedT1 | 2026-04-16SimpleBench Leaderboard |
| ProofBench | 54accuracy (%) | unverifiedT1 | 2026-04-16https://www.vals.ai/benchmarks/proof_bench |
| SimpleQA Verified | 50.6accuracy (%) | reproducedT1 | 2026-04-16Epoch AI |
| GSO-Bench | 44.12accuracy (%) | unverifiedT1 | 2026-04-16GSO Leaderboard |
| FrontierCode | 38.5accuracy (%) | unverifiedT1 | 2026-04-16https://cognition.com/frontiercode |
| APEX-Agents | 33.9accuracy (%) | unverifiedT2 | 2026-04-16 |
| HLE | 32.98accuracy (%) | unverifiedT2 | 2026-04-16 |
| FrontierMath-Tier-4-v2-Private | 31.71accuracy (%) | reproducedT1 | 2026-04-16Epoch AI |
| MirrorCode | 31.11accuracy (%) | reproducedT1 | 2026-04-16Epoch AI |
| PostTrainBench | 28.56accuracy (%) | unverifiedT2 | 2026-04-16 |
| ExploitBench | 26.5accuracy (%) | unverifiedT2 | 2026-04-16 |
| Chess Puzzles | 26.35accuracy (%) | reproducedT1 | 2026-04-16Epoch AI |
| Mystery Game Puzzles | 20.69accuracy (%) | reproducedT1 | 2026-04-16Epoch AI |
| EBR-bench | 19.05accuracy (%) | reproducedT1 | 2026-04-16Epoch AI |
| OSWorld 2.0 | 18.2accuracy (%) | unverifiedT1 | 2026-04-16https://osworld-v2.xlang.ai/ |
| CritPt | 12accuracy (%) | unverifiedT2 | 2026-04-16 |
Claims drawn from cited facts, not live model generation.
The model is scored on 24 tracked benchmarks and currently holds the top score on 1 of them. Because the results are independently reproduced, these scores are backed by outside evaluation and can be treated as verified.
3 cited facts
It trails the leader by 18.75 points on average across the disclosed benchmark suite, a wide gap that places it well behind the front-of-pack. Although it holds the top score on 1 benchmark, that single lead does not offset the overall average gap.
3 cited facts
6 facts cross-checked across data sources: 3 corroborated, 2 single-source, 1 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: curated_capability_claims, epoch_benchmarks, litellm_prices, modelsdev_models +1 more
Claude Opus 4.7 is an AI model developed by Anthropic, released 2026-04-16. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Claude Opus 4.7 has recorded scores on 24 benchmarks, each shown with its evidence status.
Claude Opus 4.7 has recorded scores on 24 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleBench, SimpleQA Verified, and 18 more. The full table above shows each score with its evidence status.
10 of 24 recorded scores (42%) are independently reproduced rather than self-reported by the lab.
Claude Opus 4.7 has 24 tracked claims: 10 independently reproduced, 14 unverified.
Listed API pricing: $5 per million input tokens, $25 per million output tokens. See the pricing block for the full breakdown.