Anthropic·released 2025-05-221 source
Claude Sonnet 4 benchmark scores: 23 benchmarks tracked. 22% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 23 benchmarks·best result 84.37 on MATH level 5 (reproduced)·5 of 23 independently reproduced·$3/$15 per M tokens
Consensus: LiteLLM · Cross-check: OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MATH level 5 | 84.37accuracy (%) | reproduced· optimizedT1 | 2025-05-22Epoch AI |
| Lech Mazur Writing | 81.4accuracy (%) | unverifiedT1 | 2025-05-22lechmazur/writing Github repository |
| GPQA diamond | 72.25accuracy (%) | reproduced· optimizedT1 | 2025-05-22Epoch AI |
| OTIS Mock AIME 2024-2025 | 71.08accuracy (%) | reproduced· optimizedT1 | 2025-05-22Epoch AI |
| METR Time Horizons | 61.97accuracy (%) | unverifiedT1 | 2025-05-22METR - Measuring AI Ability to Complete Long Tasks |
| Aider polyglot | 61.3accuracy (%) | unverified· optimizedT1 | 2025-05-22Aider LLM Leaderboards |
| DeepResearch Bench | 47.8accuracy (%) | unverifiedT1 | 2025-05-22DeepResearchBench Leaderboard |
| Fiction.LiveBench | 46.9accuracy (%) | unverifiedT1 | 2025-05-22Fiction.live leaderboard |
| WeirdML | 46.11accuracy (%) | unverifiedT1 | 2025-05-22WeirdML Leaderboard |
| OSWorld | 43.9accuracy (%) | unverified· optimizedT1 | 2025-05-22OS World Website |
| ARC-AGI | 40accuracy (%) | unverified· optimizedT1 | 2025-05-22ARC Prize Leaderboard |
| GeoBench | 37accuracy (%) | unverifiedT1 | 2025-05-22GeoBench leaderboard |
| Cybench | 35accuracy (%) | unverified· optimizedT1 | 2025-05-22Cybench leaderboard |
| SimpleBench | 34.6accuracy (%) | unverifiedT1 | 2025-05-22SimpleBench Leaderboard |
| The Agent Company | 33.1accuracy (%) | unverifiedT1 | 2025-05-22TheAgentCompany leaderboard |
| APEX-Agents | 9.3accuracy (%) | unverifiedT2 | 2025-05-22 |
| FrontierMath-2025-02-28-Private | 7.26accuracy (%) | reproduced· optimizedT1 | 2025-05-22Epoch AI |
| ARC-AGI-2 | 5.93accuracy (%) | unverifiedT2 | 2025-05-22 |
| GSO-Bench | 4.9accuracy (%) | unverifiedT1 | 2025-05-22GSO Leaderboard |
| HLE | 3.11accuracy (%) | unverifiedT2 | 2025-05-22 |
| VPCT | 1accuracy (%) | unverifiedT1 | 2025-05-22VPCT leaderboard |
| CritPt | 0.29accuracy (%) | unverifiedT2 | 2025-05-22 |
| FrontierMath-Tier-4-2025-07-01-Private | 0accuracy (%) | reproducedT1 | 2025-05-22Epoch AI |
Claims drawn from cited facts, not live model generation.
This model is scored on 23 tracked benchmarks. Of those, it currently holds no top score. Because its scores are independently reproduced, the record is backed by outside evaluation and can be trusted as verified.
3 cited facts
With an average gap of 38.24 points to the state-of-the-art leader, this model sits in a wide-gap band — well behind the front of the pack on its benchmarks. It holds the top score on none of the benchmark leaderboards, so there is no genuine category leadership yet.
2 cited facts
6 facts cross-checked across data sources: 3 corroborated, 2 single-source, 1 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, openrouter_models
Claude Sonnet 4 is an AI model developed by Anthropic, released 2025-05-22. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Claude Sonnet 4 has recorded scores on 23 benchmarks, each shown with its evidence status.
Claude Sonnet 4 has recorded scores on 23 benchmarks — Lech Mazur Writing, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, Cybench, and 17 more. The full table above shows each score with its evidence status.
5 of 23 recorded scores (22%) are independently reproduced rather than self-reported by the lab.
Claude Sonnet 4 has 23 tracked claims: 5 independently reproduced, 18 unverified.
Listed API pricing: $3 per million input tokens, $15 per million output tokens. See the pricing block for the full breakdown.