Anthropic·released 2024-06-201 source
Claude 3.5 Sonnet benchmark scores: 11 benchmarks tracked. 45% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 11 benchmarks·best result 82 on MMLU (unverified)·5 of 11 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MMLU | 82accuracy (%) | unverified· optimizedT1 | 2024-06-20Stanford CRFM Leaderboard |
| MATH level 5 | 51.68accuracy (%) | reproduced· optimizedT1 | 2024-06-20Epoch AI |
| METR Time Horizons | 47.68accuracy (%) | unverifiedT1 | 2024-06-20METR - Measuring AI Ability to Complete Long Tasks |
| GPQA diamond | 38.72accuracy (%) | reproduced· optimizedT1 | 2024-06-20Epoch AI |
| WeirdML | 30.97accuracy (%) | unverifiedT1 | 2024-06-20WeirdML Leaderboard |
| Cybench | 17.5accuracy (%) | unverified· optimizedT1 | 2024-06-20Cybench leaderboard |
| SimpleBench | 13accuracy (%) | unverifiedT1 | 2024-06-20SimpleBench Leaderboard |
| OTIS Mock AIME 2024-2025 | 6.43accuracy (%) | reproduced· optimizedT1 | 2024-06-20Epoch AI |
| FrontierMath-2025-02-28-Private | 1.81accuracy (%) | reproduced· optimizedT1 | 2024-06-20Epoch AI |
| FrontierMath-Tier-4-2025-07-01-Private | 0accuracy (%) | reproducedT1 | 2024-06-20Epoch AI |
| VPCT | 0accuracy (%) | unverifiedT1 | 2024-06-20VPCT leaderboard |
Claims drawn from cited facts, not live model generation.
This model was developed by Anthropic and released in 2024, making it a recent addition to the developer's lineup. Given its 2024 release, this model reflects Anthropic's current-generation capabilities.
4 cited facts
This model is evaluated across 11 tracked benchmarks, and among those it currently holds no top score. Because the scores have been independently reproduced by outside evaluation, the record can be trusted as verified rather than vendor-claimed.
3 cited facts
With an average score gap of 55.8 points on the disclosed harness, this model sits well behind the leaders — a wide gap, not a frontier-level result. It holds no top scores on the evaluated benchmarks, indicating no category leadership yet.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Claude 3.5 Sonnet is an AI model developed by Anthropic, released 2024-06-20. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Claude 3.5 Sonnet has recorded scores on 11 benchmarks, each shown with its evidence status.
Claude 3.5 Sonnet has recorded scores on 11 benchmarks — GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, MMLU, Cybench, and 5 more. The full table above shows each score with its evidence status.
5 of 11 recorded scores (45%) are independently reproduced rather than self-reported by the lab.
Claude 3.5 Sonnet has 11 tracked claims: 5 independently reproduced, 6 unverified.