Anthropic·released 2024-03-041 source
Claude 3 Sonnet benchmark scores: 6 benchmarks tracked. 50% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 6 benchmarks·best result 67.87 on MMLU (unverified)·3 of 6 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MMLU | 67.87accuracy (%) | unverified· optimizedT1 | 2024-03-04Stanford CRFM Leaderboard |
| Winogrande | 50.2accuracy (%) | unverified· optimizedT1 | 2024-03-04The Claude 3 Model Family: Opus, Sonnet, Haiku |
| GPQA diamond | 20.79accuracy (%) | reproduced· optimizedT1 | 2024-03-04Epoch AI |
| MATH level 5 | 18.17accuracy (%) | reproduced· optimizedT1 | 2024-03-04Epoch AI |
| WeirdML | 10.16accuracy (%) | unverifiedT1 | 2024-03-04WeirdML Leaderboard |
| OTIS Mock AIME 2024-2025 | 2.4accuracy (%) | reproduced· optimizedT1 | 2024-03-04Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by Anthropic and released in 2024.
2 cited facts
With 6 benchmarks in and no leads, the record is modest and unremarkable — which is itself information: nothing here flags either a standout strength or an outlier. Because the scores have been independently reproduced, readers can trust the reported numbers.
3 cited facts
This model holds the top score on none of the benchmarks and trails the state-of-the-art by an average of 61.95 points, placing it well behind the leaders.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Claude 3 Sonnet is an AI model developed by Anthropic, released 2024-03-04. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Claude 3 Sonnet has recorded scores on 6 benchmarks, each shown with its evidence status.
Claude 3 Sonnet has recorded scores on 6 benchmarks — Winogrande, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, MMLU. The full table above shows each score with its evidence status.
3 of 6 recorded scores (50%) are independently reproduced rather than self-reported by the lab.
Claude 3 Sonnet has 6 tracked claims: 3 independently reproduced, 3 unverified.