Anthropic·released 2024-10-221 source
Claude 3.5 Haiku benchmark scores: 13 benchmarks tracked. 38% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 13 benchmarks·best result 73.5 on Lech Mazur Writing (unverified)·5 of 13 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| Lech Mazur Writing | 73.5accuracy (%) | unverifiedT1 | 2024-10-22lechmazur/writing Github repository |
| MMLU | 65.73accuracy (%) | unverified· optimizedT1 | 2024-10-22Stanford CRFM Leaderboard |
| MATH level 5 | 46.36accuracy (%) | reproduced· optimizedT1 | 2024-10-22Epoch AI |
| GeoBench | 34accuracy (%) | unverifiedT1 | 2024-10-22GeoBench leaderboard |
| CadEval | 32accuracy (%) | unverifiedT1 | 2024-10-22CadEval Dashboard |
| WeirdML | 30.73accuracy (%) | unverifiedT1 | 2024-10-22WeirdML Leaderboard |
| Aider polyglot | 28accuracy (%) | unverified· optimizedT1 | 2024-10-22Aider LLM Leaderboards |
| Balrog | 19.3accuracy (%) | unverifiedT1 | 2024-10-22Balrog Leaderboard |
| GPQA diamond | 17.51accuracy (%) | reproduced· optimizedT1 | 2024-10-22Epoch AI |
| SimpleQA Verified | 6.7accuracy (%) | reproducedT1 | 2024-10-22Epoch AI |
| OTIS Mock AIME 2024-2025 | 4.21accuracy (%) | reproduced· optimizedT1 | 2024-10-22Epoch AI |
| FrontierMath-2025-02-28-Private | 0.6accuracy (%) | reproduced· optimizedT1 | 2024-10-22Epoch AI |
| CritPt | 0accuracy (%) | unverifiedT2 | 2024-10-22 |
Claims drawn from cited facts, not live model generation.
This model was developed by Anthropic and released in 2024.
2 cited facts
At 13 tracked benchmarks, coverage is substantial enough that the record reflects a real evaluation footprint rather than a couple of favorable runs. Topping none of them is the common case — leadership concentrates in a few entrants at any time — so the informative part of this page is where its scores sit, not the missing first place. Because its scores are independently reproduced, outside evaluation backs the record, so you can trust the numbers.
3 cited facts
On average, this model trails the state-of-the-art leader by 52.0 points per benchmark, indicating a wide gap that places it well behind the front of the pack. It leads nothing here, and the count flattens every near-miss into the same result — how close it runs to the front is the question this figure cannot answer.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Claude 3.5 Haiku is an AI model developed by Anthropic, released 2024-10-22. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Claude 3.5 Haiku has recorded scores on 13 benchmarks, each shown with its evidence status.
Claude 3.5 Haiku has recorded scores on 13 benchmarks — Lech Mazur Writing, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, CadEval, and 7 more. The full table above shows each score with its evidence status.
5 of 13 recorded scores (38%) are independently reproduced rather than self-reported by the lab.
Claude 3.5 Haiku has 13 tracked claims: 5 independently reproduced, 8 unverified.