Anthropic·released 2026-06-30✓ 2 sources
Claude Sonnet 5 benchmark scores: 16 benchmarks tracked. 44% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 16 benchmarks·best result 94.72 on OTIS Mock AIME 2024-2025 (reproduced)·7 of 16 independently reproduced·$2/$10 per M tokens
Consensus: LiteLLM · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 94.72accuracy (%) | reproduced· optimizedT1 | 2026-06-30Epoch AI |
| GPQA diamond | 87.37accuracy (%) | reproduced· optimizedT1 | 2026-06-30Epoch AI |
| ProofBench | 77accuracy (%) | unverifiedT1 | 2026-06-30https://www.vals.ai/benchmarks/proof_bench |
| WeirdML | 68.78accuracy (%) | unverifiedT1 | 2026-06-30https://htihle.github.io/weirdml.html |
| FrontierMath-Tiers-1-3-v2-Private | 65.61accuracy (%) | reproducedT1 | 2026-06-30Epoch AI |
| Surface Evolver Bench | 60accuracy (%) | unverifiedT1 | 2026-06-30https://yhenon.github.io/surface-evolver-llm-eval/ |
| DeepSWE | 53.85accuracy (%) | unverifiedT1 | 2026-06-30https://deepswe.datacurve.ai/ |
| SimpleBench | 52.72accuracy (%) | unverifiedT1 | 2026-06-30SimpleBench Leaderboard |
| FrontierCode | 42.7accuracy (%) | unverifiedT1 | 2026-06-30https://cognition.com/frontiercode |
| GSO-Bench | 37.25accuracy (%) | unverifiedT1 | 2026-06-30https://gso-bench.github.io/index.html |
| APEX-Agents | 32.5accuracy (%) | unverifiedT2 | 2026-06-30 |
| Chess Puzzles | 31.61accuracy (%) | reproducedT1 | 2026-06-30Epoch AI |
| FrontierMath-Tier-4-v2-Private | 29.27accuracy (%) | reproducedT1 | 2026-06-30Epoch AI |
| Mystery Game Puzzles | 28.4accuracy (%) | reproducedT1 | 2026-06-30Epoch AI |
| SimpleQA Verified | 25accuracy (%) | reproducedT1 | 2026-06-30Epoch AI |
| CritPt | 16.86accuracy (%) | unverifiedT2 | 2026-06-30 |
Claims drawn from cited facts, not live model generation.
Developed by Anthropic, this model was released in 2026.
2 cited facts
This model is tracked across 16 benchmarks, and it currently holds no top score on any of them. Because the record is independently reproduced, outside evaluation backs these scores, so the benchmark results can be read as verified.
3 cited facts
Trailing the leader by an average of 23.52 points on the tracked benchmarks, this model sits well behind the frontier rather than competing at the top. It holds the top score on none of the benchmarks, so there is no category leadership yet.
2 cited facts
6 facts cross-checked across data sources: 1 corroborated, 5 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: curated_models, epoch_benchmarks, litellm_prices
Claude Sonnet 5 is an AI model developed by Anthropic, released 2026-06-30. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Claude Sonnet 5 has recorded scores on 16 benchmarks, each shown with its evidence status.
Claude Sonnet 5 has recorded scores on 16 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleBench, SimpleQA Verified, and 10 more. The full table above shows each score with its evidence status.
7 of 16 recorded scores (44%) are independently reproduced rather than self-reported by the lab.
Claude Sonnet 5 has 16 tracked claims: 7 independently reproduced, 9 unverified.
Listed API pricing: $2 per million input tokens, $10 per million output tokens. See the pricing block for the full breakdown.