Anthropic·released 2023-11-211 source
Claude 2.1 benchmark scores: 4 benchmarks tracked. 50% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 4 benchmarks·best result 64.67 on MMLU (unverified)·2 of 4 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MMLU | 64.67accuracy (%) | unverified· optimizedT1 | 2023-11-21Stanford CRFM Leaderboard |
| GPQA diamond | 10.61accuracy (%) | reproduced· optimizedT1 | 2023-11-21Epoch AI |
| WeirdML | 7.06accuracy (%) | unverifiedT1 | 2023-11-21WeirdML Leaderboard |
| OTIS Mock AIME 2024-2025 | 1.85accuracy (%) | reproduced· optimizedT1 | 2023-11-21Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by Anthropic and released in 2023.
2 cited facts
Four tracked results make a thin file — enough to locate the model, not enough to characterize it. Not leading any of its benchmarks places it mid-field; on this site that reads as a position marker, not a strike against the model. Because its scores have been independently reproduced by external evaluators, this record can be trusted.
3 cited facts
The model trails the state-of-the-art leader by an average of 70.15 points, a wide gap that places it well behind the front-of-pack systems.
1 cited fact
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Claude 2.1 is an AI model developed by Anthropic, released 2023-11-21. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Claude 2.1 has recorded scores on 4 benchmarks, each shown with its evidence status.
Claude 2.1 has recorded scores on 4 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, MMLU. The full table above shows each score with its evidence status.
2 of 4 recorded scores (50%) are independently reproduced rather than self-reported by the lab.
Claude 2.1 has 4 tracked claims: 2 independently reproduced, 2 unverified.