Anthropic·released 2026-07-241 source
Claude Opus 5 benchmark scores: 17 benchmarks tracked, leading 4. 47% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 4 of 17 benchmarks·best result 99 on ProofBench (unverified)·8 of 17 independently reproduced
Head-to-headClaude Opus 5 vs GPT-5.6 SolLeads on: DeepSWE, EBR-bench, Mystery Game Puzzles, ProofBench
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| ProofBench | 99accuracy (%) | unverifiedT1 | 2026-07-24https://www.vals.ai/benchmarks/proof_bench |
| OTIS Mock AIME 2024-2025 | 98.89accuracy (%) | reproduced· optimizedT1 | 2026-07-24Epoch AI |
| ARC-AGI | 97.5accuracy (%) | unverified· optimizedT1 | 2026-07-24https://arcprize.org/leaderboard |
| GPQA diamond | 91.84accuracy (%) | reproduced· optimizedT1 | 2026-07-24Epoch AI |
| WeirdML | 91.78accuracy (%) | unverifiedT1 | 2026-07-24https://htihle.github.io/weirdml.html |
| ARC-AGI-2 | 90.42accuracy (%) | unverifiedT2 | 2026-07-24 |
| FrontierMath-Tiers-1-3-v2-Private | 85.61accuracy (%) | reproducedT1 | 2026-07-24Epoch AI |
| DeepSWE | 73.65accuracy (%) | unverifiedT1 | 2026-07-24https://deepswe.datacurve.ai/ |
| FrontierMath-Tier-4-v2-Private | 73.17accuracy (%) | reproducedT1 | 2026-07-24Epoch AI |
| SimpleQA Verified | 56.7accuracy (%) | reproducedT1 | 2026-07-24Epoch AI |
| Mystery Game Puzzles | 54.84accuracy (%) | reproducedT1 | 2026-07-24Epoch AI |
| FrontierCode | 53.4accuracy (%) | unverifiedT1 | 2026-07-24https://cognition.com/frontiercode |
| EBR-bench | 50accuracy (%) | reproducedT1 | 2026-07-24Epoch AI |
| APEX-Agents | 43.5accuracy (%) | unverifiedT2 | 2026-07-24 |
| Chess Puzzles | 38.97accuracy (%) | reproducedT1 | 2026-07-24Epoch AI |
| PostTrainBench | 34.06accuracy (%) | unverifiedT2 | 2026-07-24 |
| CritPt | 29.14accuracy (%) | unverifiedT2 | 2026-07-24 |
Claims drawn from cited facts, not live model generation.
This model, developed by Anthropic, was released in 2026.
2 cited facts
This model is tracked across 17 benchmarks, and it currently holds the top score on 4 of them. Because these results have been independently reproduced by outside evaluation, they should be read as verified rather than as mere vendor claims.
3 cited facts
It trails the top benchmark score by an average of 4.71 points, a moderate gap that reads as competitive but not front-of-pack. It nevertheless holds the best score on 4 benchmarks, indicating genuine category leadership in specific settings.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Claude Opus 5 is an AI model developed by Anthropic, released 2026-07-24. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Claude Opus 5 has recorded scores on 17 benchmarks, each shown with its evidence status.
Claude Opus 5 has recorded scores on 17 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleQA Verified, CritPt, and 11 more. The full table above shows each score with its evidence status.
8 of 17 recorded scores (47%) are independently reproduced rather than self-reported by the lab.
Claude Opus 5 has 17 tracked claims: 8 independently reproduced, 9 unverified.