Anthropic·released 2025-11-241 source
Claude Opus 4.5 benchmark scores: 28 benchmarks tracked. 32% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 28 benchmarks·best result 86.1 on OTIS Mock AIME 2024-2025 (reproduced)·9 of 28 independently reproduced·$5/$25 per M tokens
Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 86.1accuracy (%) | reproduced· optimizedT1 | 2025-11-24Epoch AI |
| Cybench | 82accuracy (%) | self-reported· optimizedT1 | 2025-11-24Opus 4.5 System Card |
| GPQA diamond | 81.4accuracy (%) | reproduced· optimizedT1 | 2025-11-24Epoch AI |
| ARC-AGI | 80accuracy (%) | unverified· optimizedT1 | 2025-11-24ARC Prize Leaderboard |
| SWE-Bench verified | 76.65accuracy (%) | reproduced· optimizedT1 | 2025-11-24Epoch AI |
| GeoBench | 75accuracy (%) | unverifiedT1 | 2025-11-24GeoBench leaderboard |
| METR Time Horizons | 74.97accuracy (%) | unverifiedT1 | 2025-11-24METR - Measuring AI Ability to Complete Long Tasks |
| OSWorld | 66.3accuracy (%) | unverified· optimizedT1 | 2025-11-24Anthropic annoucement |
| WeirdML | 63.7accuracy (%) | unverifiedT1 | 2025-11-24WeirdML Leaderboard |
| Terminal Bench | 63.1accuracy (%) | unverified· optimizedT1 | 2025-11-24Terminal-Bench v2 Leaderboard |
| SimpleBench | 54.4accuracy (%) | unverifiedT1 | 2025-11-24SimpleBench Leaderboard |
| GDPval | 45.5accuracy (%) | unverifiedT2 | 2025-11-24 |
| Balrog | 43.5accuracy (%) | unverifiedT1 | 2025-11-24Balrog Leaderboard |
| SimpleQA Verified | 41.8accuracy (%) | reproducedT1 | 2025-11-24Epoch AI |
| ARC-AGI-2 | 37.64accuracy (%) | unverifiedT2 | 2025-11-24 |
| ProofBench | 36accuracy (%) | unverifiedT1 | 2025-11-24https://www.vals.ai/benchmarks/proof_bench |
| FrontierMath-Tiers-1-3-v2-Private | 34.39accuracy (%) | reproducedT1 | 2025-11-24Epoch AI |
| GSO-Bench | 26.5accuracy (%) | unverifiedT1 | 2025-11-24GSO Leaderboard |
| HLE | 21.43accuracy (%) | unverifiedT2 | 2025-11-24 |
| CL-bench | 21.1accuracy (%) | unverifiedT2 | 2025-11-24 |
| APEX-Agents | 20.7accuracy (%) | unverifiedT2 | 2025-11-24 |
| PostTrainBench | 17.29accuracy (%) | unverifiedT2 | 2025-11-24 |
| EBR-bench | 14.29accuracy (%) | reproducedT1 | 2025-11-24Epoch AI |
| Mystery Game Puzzles | 14.08accuracy (%) | reproducedT1 | 2025-11-24Epoch AI |
| VPCT | 10accuracy (%) | unverifiedT1 | 2025-11-24Tweet from VPCT creator |
| Chess Puzzles | 7.41accuracy (%) | reproducedT1 | 2025-11-24Epoch AI |
| FrontierMath-Tier-4-v2-Private | 4.88accuracy (%) | reproducedT1 | 2025-11-24Epoch AI |
| Remote Labor Index | 3.75accuracy (%) | unverifiedT2 | 2025-11-24 |
Claims drawn from cited facts, not live model generation.
This model was developed by Anthropic and released in 2025, making it part of that developer's ongoing model line. Since no parameter count is provided, this model's scale is unspecified, and no size-related trade-off is stated.
2 cited facts
This model is tracked across 28 benchmarks and currently holds no top score on any of them. Because the record has been independently reproduced, these results are backed by outside evaluation and can be treated as verified.
3 cited facts
The model trails the SOTA leader by an average of 28.09 points under the disclosed harness, a wide gap placing it well behind the leaders. It holds the top score on none of the tracked benchmarks, so it has no genuine category leadership yet.
2 cited facts
6 facts cross-checked across data sources: 3 corroborated, 2 single-source, 1 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, modelsdev_models, openrouter_models
Claude Opus 4.5 is an AI model developed by Anthropic, released 2025-11-24. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Claude Opus 4.5 has recorded scores on 28 benchmarks, each shown with its evidence status.
Claude Opus 4.5 has recorded scores on 28 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, Cybench, SimpleBench, and 22 more. The full table above shows each score with its evidence status.
9 of 28 recorded scores (32%) are independently reproduced rather than self-reported by the lab.
Claude Opus 4.5 has 28 tracked claims: 9 independently reproduced, 1 self-reported, 18 unverified.
Listed API pricing: $5 per million input tokens, $25 per million output tokens. See the pricing block for the full breakdown.