Anthropic·released 2024-03-041 source
Claude 3 Opus benchmark scores: 10 benchmarks tracked. 40% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 10 benchmarks·best result 79.47 on MMLU (unverified)·4 of 10 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MMLU | 79.47accuracy (%) | unverified· optimizedT1 | 2024-03-04Stanford CRFM Leaderboard |
| Winogrande | 77accuracy (%) | unverified· optimizedT1 | 2024-03-04The Claude 3 Model Family: Opus, Sonnet, Haiku |
| METR Time Horizons | 37.75accuracy (%) | unverifiedT1 | 2024-03-04METR - Measuring AI Ability to Complete Long Tasks |
| MATH level 5 | 37.48accuracy (%) | reproduced· optimizedT1 | 2024-03-04Epoch AI |
| GPQA diamond | 29.55accuracy (%) | reproduced· optimizedT1 | 2024-03-04Epoch AI |
| WeirdML | 23.18accuracy (%) | unverifiedT1 | 2024-03-04WeirdML Leaderboard |
| Cybench | 10accuracy (%) | unverified· optimizedT1 | 2024-03-04Cybench leaderboard |
| SimpleBench | 8.2accuracy (%) | unverifiedT1 | 2024-03-04SimpleBench Leaderboard |
| OTIS Mock AIME 2024-2025 | 4.63accuracy (%) | reproduced· optimizedT1 | 2024-03-04Epoch AI |
| Chess Puzzles | 0.04accuracy (%) | reproducedT1 | 2024-03-04Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by Anthropic and released in 2024.
2 cited facts
This model is scored on 10 tracked benchmarks. It currently holds no top score on any of those benchmarks. Because the results are independently reproduced, outside evaluation backs these scores, so the record can be treated as verified.
3 cited facts
With an average gap of 55.07 points to the top score, this model sits well behind the leaders on its benchmarks. It does not hold the top score on any benchmark, so it has no current category leadership.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Claude 3 Opus is an AI model developed by Anthropic, released 2024-03-04. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Claude 3 Opus has recorded scores on 10 benchmarks, each shown with its evidence status.
Claude 3 Opus has recorded scores on 10 benchmarks — Winogrande, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, MMLU, and 4 more. The full table above shows each score with its evidence status.
4 of 10 recorded scores (40%) are independently reproduced rather than self-reported by the lab.
Claude 3 Opus has 10 tracked claims: 4 independently reproduced, 6 unverified.