Anthropic·released 2024-03-041 source
Claude 3 Haiku benchmark scores: 8 benchmarks tracked. 38% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 8 benchmarks·best result 65.07 on MMLU (unverified)·3 of 8 independently reproduced·$0.25/$1.25 per M tokens
Consensus: LiteLLM · Cross-check: OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MMLU | 65.07accuracy (%) | unverified· optimizedT1 | 2024-03-04Stanford CRFM Leaderboard |
| ScienceQA | 62.67accuracy (%) | self-reported· optimizedT1 | 2024-03-04Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| Winogrande | 48.4accuracy (%) | unverified· optimizedT1 | 2024-03-04The Claude 3 Model Family: Opus, Sonnet, Haiku |
| GPQA diamond | 15.07accuracy (%) | reproduced· optimizedT1 | 2024-03-04Epoch AI |
| MATH level 5 | 14.88accuracy (%) | reproduced· optimizedT1 | 2024-03-04Epoch AI |
| CadEval | 12accuracy (%) | unverifiedT1 | 2024-03-04CadEval Dashboard |
| WeirdML | 9.84accuracy (%) | unverifiedT1 | 2024-03-04WeirdML Leaderboard |
| OTIS Mock AIME 2024-2025 | 1.71accuracy (%) | reproduced· optimizedT1 | 2024-03-04Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by Anthropic and released in 2024.
2 cited facts
Mid-sized and unled: 8 benchmarks scored, none currently topped — enough data to form a real profile, with the missing first place a standing rather than a flaw. Because the scores are independently reproduced, outside evaluation backs up these results, so the record can be trusted as verified.
3 cited facts
With an average gap of 58.79 points to the state-of-the-art, this model is well behind the leaders, and it does not hold a top score on any benchmark, confirming it is not at the frontier.
2 cited facts
6 facts cross-checked across data sources: 3 corroborated, 2 single-source, 1 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, openrouter_models
Claude 3 Haiku is an AI model developed by Anthropic, released 2024-03-04. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Claude 3 Haiku has recorded scores on 8 benchmarks, each shown with its evidence status.
Claude 3 Haiku has recorded scores on 8 benchmarks — Winogrande, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, CadEval, and 2 more. The full table above shows each score with its evidence status.
3 of 8 recorded scores (38%) are independently reproduced rather than self-reported by the lab.
Claude 3 Haiku has 8 tracked claims: 3 independently reproduced, 1 self-reported, 4 unverified.
Listed API pricing: $0.25 per million input tokens, $1.25 per million output tokens. See the pricing block for the full breakdown.