Anthropic·released 2025-10-151 source
Claude Haiku 4.5 benchmark scores: 15 benchmarks tracked. 47% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 15 benchmarks·best result 96.36 on MATH level 5 (reproduced)·7 of 15 independently reproduced·$1/$5 per M tokens
Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MATH level 5 | 96.36accuracy (%) | reproduced· optimizedT1 | 2025-10-15Epoch AI |
| OTIS Mock AIME 2024-2025 | 66.63accuracy (%) | reproduced· optimizedT1 | 2025-10-15Epoch AI |
| GPQA diamond | 61.62accuracy (%) | reproduced· optimizedT1 | 2025-10-15Epoch AI |
| ARC-AGI | 47.67accuracy (%) | unverified· optimizedT2 | 2025-10-15 |
| WeirdML | 45.4accuracy (%) | unverifiedT1 | 2025-10-15WeirdML Leaderboard |
| Terminal Bench | 35.5accuracy (%) | unverified· optimizedT1 | 2025-10-15Terminal-Bench v2 Leaderboard |
| Balrog | 31.2accuracy (%) | unverifiedT1 | 2025-10-15Balrog Leaderboard |
| ExploitBench | 13.7accuracy (%) | unverifiedT2 | 2025-10-15 |
| FrontierMath-2025-02-28-Private | 10.36accuracy (%) | reproduced· optimizedT1 | 2025-10-15Epoch AI |
| APEX-Agents | 8.9accuracy (%) | unverifiedT2 | 2025-10-15 |
| SimpleQA Verified | 5.9accuracy (%) | reproducedT1 | 2025-10-15Epoch AI |
| ARC-AGI-2 | 4.03accuracy (%) | unverifiedT2 | 2025-10-15 |
| FrontierMath-Tier-4-2025-07-01-Private | 3.47accuracy (%) | reproducedT1 | 2025-10-15Epoch AI |
| Chess Puzzles | 3.2accuracy (%) | reproducedT1 | 2025-10-15Epoch AI |
| CritPt | 0accuracy (%) | unverifiedT2 | 2025-10-15 |
Claims drawn from cited facts, not live model generation.
This model was developed by Anthropic. It was released in 2025.
2 cited facts
This model is tracked on 15 benchmarks and currently holds no top score on any of them. Its benchmark record has been independently reproduced, so the reported results can be trusted as verified.
3 cited facts
It trails the leader by 43.12 points on average, a wide gap that places it well behind the front of the pack. It holds the top score on none of the tracked benchmarks, so it has no category leadership yet.
2 cited facts
6 facts cross-checked across data sources: 3 corroborated, 2 single-source, 1 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, modelsdev_models, openrouter_models
Claude Haiku 4.5 is an AI model developed by Anthropic, released 2025-10-15. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Claude Haiku 4.5 has recorded scores on 15 benchmarks, each shown with its evidence status.
Claude Haiku 4.5 has recorded scores on 15 benchmarks — GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, FrontierMath-2025-02-28-Private, and 9 more. The full table above shows each score with its evidence status.
7 of 15 recorded scores (47%) are independently reproduced rather than self-reported by the lab.
Claude Haiku 4.5 has 15 tracked claims: 7 independently reproduced, 8 unverified.
Listed API pricing: $1 per million input tokens, $5 per million output tokens. See the pricing block for the full breakdown.