Anthropic·released 2023-07-111 source
Claude 2 benchmark scores: 5 benchmarks tracked. 60% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 5 benchmarks·best result 87.5 on TriviaQA (self-reported)·3 of 5 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| TriviaQA | 87.5accuracy (%) | self-reported· optimizedT1 | 2023-07-11Model Card and Evaluations for Claude Models |
| MMLU | 71.33accuracy (%) | self-reported· optimizedT1 | 2023-07-11Model Card and Evaluations for Claude Models |
| GPQA diamond | 12.88accuracy (%) | reproduced· optimizedT1 | 2023-07-11Epoch AI |
| MATH level 5 | 11.73accuracy (%) | reproduced· optimizedT1 | 2023-07-11Epoch AI |
| OTIS Mock AIME 2024-2025 | 2.4accuracy (%) | reproduced· optimizedT1 | 2023-07-11Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by Anthropic and released in 2023.
2 cited facts
Only 5 benchmarks carry scores for it, so treat any pattern here as a sketch — a base this small can flip on a single new result. Holding no top score is a position, not a grade: it marks the model as in the field rather than ahead of it. Because the scores have been independently reproduced, the record can be trusted.
3 cited facts
This model trails the state-of-the-art leader by an average of 55.36 points, a wide gap. It holds the top score on none of the tracked benchmarks, confirming it is not a leader in any category.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Claude 2 is an AI model developed by Anthropic, released 2023-07-11. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Claude 2 has recorded scores on 5 benchmarks, each shown with its evidence status.
Claude 2 has recorded scores on 5 benchmarks — TriviaQA, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, MMLU. The full table above shows each score with its evidence status.
3 of 5 recorded scores (60%) are independently reproduced rather than self-reported by the lab.
Claude 2 has 5 tracked claims: 3 independently reproduced, 2 self-reported.