Anthropic·released 2025-09-291 source
Claude Sonnet 4.5 benchmark scores: 27 benchmarks tracked. 37% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 27 benchmarks·best result 97.73 on MATH level 5 (reproduced)·10 of 27 independently reproduced·$3/$15 per M tokens
Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MATH level 5 | 97.73accuracy (%) | reproduced· optimizedT1 | 2025-09-29Epoch AI |
| OTIS Mock AIME 2024-2025 | 77.76accuracy (%) | reproduced· optimizedT1 | 2025-09-29Epoch AI |
| GPQA diamond | 76.43accuracy (%) | reproduced· optimizedT1 | 2025-09-29Epoch AI |
| SWE-Bench verified | 71.28accuracy (%) | reproduced· optimizedT1 | 2025-09-29Epoch AI |
| METR Time Horizons | 67.38accuracy (%) | unverifiedT1 | 2025-09-29METR - Measuring AI Ability to Complete Long Tasks |
| ARC-AGI | 63.67accuracy (%) | unverified· optimizedT2 | 2025-09-29 |
| OSWorld | 62.9accuracy (%) | unverified· optimizedT1 | 2025-09-29OS World Website |
| Cybench | 60accuracy (%) | self-reported· optimizedT1 | 2025-09-29Opus 4.5 System Card |
| DeepResearch Bench | 52.6accuracy (%) | unverifiedT1 | 2025-09-29DeepResearchBench Leaderboard |
| WeirdML | 47.71accuracy (%) | unverifiedT1 | 2025-09-29WeirdML Leaderboard |
| Terminal Bench | 46.5accuracy (%) | unverified· optimizedT1 | 2025-09-29Terminal-Bench v2 Leaderboard |
| SimpleBench | 45.16accuracy (%) | unverifiedT1 | 2025-09-29SimpleBench Leaderboard |
| GDPval | 42.5accuracy (%) | unverifiedT2 | 2025-09-29 |
| FrontierMath-Tiers-1-3-v2-Private | 23.86accuracy (%) | reproducedT1 | 2025-09-29Epoch AI |
| SimpleQA Verified | 23.6accuracy (%) | reproducedT1 | 2025-09-29Epoch AI |
| ProofBench | 19accuracy (%) | unverifiedT1 | 2025-09-29https://www.vals.ai/benchmarks/proof_bench |
| GSO-Bench | 14.71accuracy (%) | unverifiedT1 | 2025-09-29GSO Leaderboard |
| ARC-AGI-2 | 13.61accuracy (%) | unverifiedT2 | 2025-09-29 |
| PostTrainBench | 9.94accuracy (%) | unverifiedT2 | 2025-09-29 |
| VPCT | 9.7accuracy (%) | unverifiedT1 | 2025-09-29VPCT leaderboard |
| HLE | 9.37accuracy (%) | unverifiedT2 | 2025-09-29 |
| Mystery Game Puzzles | 8.57accuracy (%) | reproducedT1 | 2025-09-29Epoch AI |
| Chess Puzzles | 7.41accuracy (%) | reproducedT1 | 2025-09-29Epoch AI |
| FrontierMath-Tier-4-v2-Private | 2.44accuracy (%) | reproducedT1 | 2025-09-29Epoch AI |
| EBR-bench | 2.38accuracy (%) | reproducedT1 | 2025-09-29Epoch AI |
| Remote Labor Index | 2.08accuracy (%) | unverifiedT2 | 2025-09-29 |
| CritPt | 1.14accuracy (%) | unverifiedT2 | 2025-09-29 |
Claims drawn from cited facts, not live model generation.
Developed by Anthropic, this model was released in 2025.
2 cited facts
This model is tracked across 27 benchmarks and currently holds no top score on any of them. Its benchmark results are independently reproduced, meaning external evaluation supports the scores and they can be treated as verified.
3 cited facts
Based on the disclosed harness, the model trails the leader by an average of 36.96 points across benchmarks, a wide gap that places it well behind the front of the pack. It holds the top score on 0 benchmarks, meaning it shows no current category leadership on the tracked leaderboards.
2 cited facts
6 facts cross-checked across data sources: 2 corroborated, 2 single-source, 2 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, modelsdev_models, openrouter_models
Claude Sonnet 4.5 is an AI model developed by Anthropic, released 2025-09-29. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Claude Sonnet 4.5 has recorded scores on 27 benchmarks, each shown with its evidence status.
Claude Sonnet 4.5 has recorded scores on 27 benchmarks — GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, Cybench, and 21 more. The full table above shows each score with its evidence status.
10 of 27 recorded scores (37%) are independently reproduced rather than self-reported by the lab.
Claude Sonnet 4.5 has 27 tracked claims: 10 independently reproduced, 1 self-reported, 16 unverified.
Listed API pricing: $3 per million input tokens, $15 per million output tokens (prices disputed across sources). See the pricing block for the full breakdown.