Anthropic·released 2024-10-221 source
Claude 3.5 Sonnet (October 2024) benchmark scores: 18 benchmarks tracked. 28% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 18 benchmarks·best result 83.07 on MMLU (unverified)·5 of 18 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MMLU | 83.07accuracy (%) | unverified· optimizedT1 | 2024-10-22Stanford CRFM Leaderboard |
| Lech Mazur Writing | 80.3accuracy (%) | unverifiedT1 | 2024-10-22lechmazur/writing Github repository |
| GeoBench | 62accuracy (%) | unverifiedT1 | 2024-10-22GeoBench leaderboard |
| MATH level 5 | 56.95accuracy (%) | reproduced· optimizedT1 | 2024-10-22Epoch AI |
| METR Time Horizons | 52.7accuracy (%) | unverifiedT1 | 2024-10-22METR - Measuring AI Ability to Complete Long Tasks |
| Aider polyglot | 51.6accuracy (%) | unverified· optimizedT1 | 2024-10-22Aider LLM Leaderboards |
| CadEval | 48accuracy (%) | unverifiedT1 | 2024-10-22CadEval Dashboard |
| GPQA diamond | 40.4accuracy (%) | reproduced· optimizedT1 | 2024-10-22Epoch AI |
| WeirdML | 39.97accuracy (%) | unverifiedT1 | 2024-10-22WeirdML Leaderboard |
| Balrog | 32.6accuracy (%) | unverifiedT1 | 2024-10-22Balrog Leaderboard |
| SimpleBench | 29.68accuracy (%) | unverifiedT1 | 2024-10-22SimpleBench Leaderboard |
| The Agent Company | 24accuracy (%) | unverifiedT1 | 2024-10-22TheAgentCompany experiment results github |
| OTIS Mock AIME 2024-2025 | 8.38accuracy (%) | reproduced· optimizedT1 | 2024-10-22Epoch AI |
| GSO-Bench | 4.6accuracy (%) | unverifiedT1 | 2024-10-22GSO Leaderboard |
| FrontierMath-2025-02-28-Private | 3.63accuracy (%) | reproduced· optimizedT1 | 2024-10-22Epoch AI |
| FrontierMath-Tier-4-2025-07-01-Private | 0accuracy (%) | reproducedT1 | 2024-10-22Epoch AI |
| VPCT | 0accuracy (%) | unverifiedT1 | 2024-10-22VPCT leaderboard |
| HLE | 0accuracy (%) | unverifiedT2 | 2024-10-22 |
Claims drawn from cited facts, not live model generation.
Developed by Anthropic, this model originates from Anthropic's research and product efforts. It was released in 2024, making it a recent addition to the AI landscape.
2 cited facts
This model is assessed across 18 tracked benchmarks. It currently holds no top score on any of them. Because the results are independently reproduced, the record can be trusted.
3 cited facts
With an average gap of 40.03 points behind the top score, this model sits well behind the frontier on its benchmarks. It also holds the top score on zero benchmarks, so there is no category leadership to point to.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Claude 3.5 Sonnet (October 2024) is an AI model developed by Anthropic, released 2024-10-22. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Claude 3.5 Sonnet (October 2024) has recorded scores on 18 benchmarks, each shown with its evidence status.
Claude 3.5 Sonnet (October 2024) has recorded scores on 18 benchmarks — Lech Mazur Writing, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, CadEval, and 12 more. The full table above shows each score with its evidence status.
5 of 18 recorded scores (28%) are independently reproduced rather than self-reported by the lab.
Claude 3.5 Sonnet (October 2024) has 18 tracked claims: 5 independently reproduced, 13 unverified.