Anthropic·released 2025-05-221 source
Claude Opus 4 benchmark scores: 21 benchmarks tracked. 29% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 21 benchmarks·best result 85.05 on MATH level 5 (reproduced)·6 of 21 independently reproduced·$15/$75 per M tokens
Consensus: LiteLLM · Cross-check: OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MATH level 5 | 85.05accuracy (%) | reproduced· optimizedT1 | 2025-05-22Epoch AI |
| Lech Mazur Writing | 83.6accuracy (%) | unverifiedT1 | 2025-05-22lechmazur/writing Github repository |
| Aider polyglot | 72accuracy (%) | unverified· optimizedT1 | 2025-05-22Aider LLM Leaderboards |
| SWE-Bench verified | 70.66accuracy (%) | reproduced· optimizedT1 | 2025-05-22Epoch AI |
| GPQA diamond | 68.35accuracy (%) | reproduced· optimizedT1 | 2025-05-22Epoch AI |
| OTIS Mock AIME 2024-2025 | 64.41accuracy (%) | reproduced· optimizedT1 | 2025-05-22Epoch AI |
| METR Time Horizons | 63.93accuracy (%) | unverifiedT1 | 2025-05-22METR - Measuring AI Ability to Complete Long Tasks |
| Fiction.LiveBench | 61.1accuracy (%) | unverifiedT1 | 2025-05-22Fiction.live leaderboard |
| SimpleBench | 50.56accuracy (%) | unverifiedT1 | 2025-05-22SimpleBench Leaderboard |
| GeoBench | 49accuracy (%) | unverifiedT1 | 2025-05-22GeoBench leaderboard |
| DeepResearch Bench | 49accuracy (%) | unverifiedT1 | 2025-05-22DeepResearchBench Leaderboard |
| WeirdML | 43.4accuracy (%) | unverifiedT1 | 2025-05-22WeirdML Leaderboard |
| Cybench | 38accuracy (%) | unverified· optimizedT1 | 2025-05-22Cybench leaderboard |
| ARC-AGI | 35.7accuracy (%) | unverified· optimizedT1 | 2025-05-22ARC Prize Leaderboard |
| ARC-AGI-2 | 8.61accuracy (%) | unverifiedT2 | 2025-05-22 |
| FrontierMath-2025-02-28-Private | 7.86accuracy (%) | reproduced· optimizedT1 | 2025-05-22Epoch AI |
| VPCT | 7accuracy (%) | unverifiedT1 | 2025-05-22VPCT leaderboard |
| FrontierMath-Tier-4-2025-07-01-Private | 6.94accuracy (%) | reproducedT1 | 2025-05-22Epoch AI |
| GSO-Bench | 6.9accuracy (%) | unverifiedT1 | 2025-05-22GSO Leaderboard |
| HLE | 6.22accuracy (%) | unverifiedT2 | 2025-05-22 |
| CritPt | 0.3accuracy (%) | unverifiedT2 | 2025-05-22 |
Claims drawn from cited facts, not live model generation.
This model was developed by Anthropic and released in 2025.
2 cited facts
This model is scored on 21 tracked benchmarks and currently holds no top score among them. Because the scores are independently reproduced, outside evaluation backs the record, so the figures can be treated as verified.
3 cited facts
With an average gap of 35.86 points behind the SOTA score, this model is well behind the leaders on its benchmark suite. It holds the top score on none of the benchmarks, meaning it currently shows no category leadership.
2 cited facts
6 facts cross-checked across data sources: 3 corroborated, 2 single-source, 1 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, openrouter_models
Claude Opus 4 is an AI model developed by Anthropic, released 2025-05-22. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Claude Opus 4 has recorded scores on 21 benchmarks, each shown with its evidence status.
Claude Opus 4 has recorded scores on 21 benchmarks — Lech Mazur Writing, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, Cybench, and 15 more. The full table above shows each score with its evidence status.
6 of 21 recorded scores (29%) are independently reproduced rather than self-reported by the lab.
Claude Opus 4 has 21 tracked claims: 6 independently reproduced, 15 unverified.
Listed API pricing: $15 per million input tokens, $75 per million output tokens. See the pricing block for the full breakdown.