Anthropic·released 2025-02-241 source
Claude 3.7 Sonnet benchmark scores: 21 benchmarks tracked. 24% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 21 benchmarks·best result 91.16 on MATH level 5 (reproduced)·5 of 21 independently reproduced·$3/$15 per M tokens
Consensus: LiteLLM · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MATH level 5 | 91.16accuracy (%) | reproduced· optimizedT1 | 2025-02-24Epoch AI |
| Fiction.LiveBench | 83.3accuracy (%) | unverifiedT1 | 2025-02-24Fiction.live leaderboard |
| Lech Mazur Writing | 81.1accuracy (%) | unverifiedT1 | 2025-02-24lechmazur/writing Github repository |
| GPQA diamond | 72.98accuracy (%) | reproduced· optimizedT1 | 2025-02-24Epoch AI |
| GeoBench | 68accuracy (%) | unverifiedT1 | 2025-02-24GeoBench leaderboard |
| Aider polyglot | 64.9accuracy (%) | unverified· optimizedT1 | 2025-02-24Aider LLM Leaderboards |
| SWE-Bench verified | 60.95accuracy (%) | reproduced· optimizedT1 | 2025-02-24Epoch AI |
| METR Time Horizons | 59.97accuracy (%) | unverifiedT1 | 2025-02-24METR - Measuring AI Ability to Complete Long Tasks |
| OTIS Mock AIME 2024-2025 | 57.74accuracy (%) | reproduced· optimizedT1 | 2025-02-24Epoch AI |
| CadEval | 54accuracy (%) | unverifiedT1 | 2025-02-24CadEval Dashboard |
| DeepResearch Bench | 43.6accuracy (%) | unverifiedT1 | 2025-02-24DeepResearchBench Leaderboard |
| OSWorld | 35.8accuracy (%) | unverified· optimizedT1 | 2025-02-24OS World Website |
| SimpleBench | 35.68accuracy (%) | unverifiedT1 | 2025-02-24SimpleBench Leaderboard |
| The Agent Company | 30.9accuracy (%) | unverifiedT1 | 2025-02-24TheAgentCompany leaderboard |
| ARC-AGI | 28.6accuracy (%) | unverified· optimizedT1 | 2025-02-24ARC Prize Leaderboard |
| Cybench | 20accuracy (%) | unverified· optimizedT1 | 2025-02-24Cybench leaderboard |
| VPCT | 8.5accuracy (%) | unverifiedT1 | 2025-02-24VPCT leaderboard |
| FrontierMath-2025-02-28-Private | 7.26accuracy (%) | reproduced· optimizedT1 | 2025-02-24Epoch AI |
| GSO-Bench | 3.8accuracy (%) | unverifiedT1 | 2025-02-24GSO Leaderboard |
| HLE | 3.4accuracy (%) | unverifiedT2 | 2025-02-24 |
| ARC-AGI-2 | 0.9accuracy (%) | unverifiedT2 | 2025-02-24 |
Claims drawn from cited facts, not live model generation.
This model originates from Anthropic and was released in 2025, reflecting the developer's ongoing work on advanced AI systems.
2 cited facts
This model is evaluated on 21 tracked benchmarks and holds no current top score on any of them. Its results are independently reproduced, so outside evaluation supports these numbers.
3 cited facts
With an average SOTA gap of 35.83 points, this model trails the leader by a wide margin and sits well behind the front of the pack on its benchmarks. It holds the top score on none of the tracked benchmarks, so it shows no category leadership yet.
2 cited facts
6 facts cross-checked across data sources: 6 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices
Claude 3.7 Sonnet is an AI model developed by Anthropic, released 2025-02-24. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Claude 3.7 Sonnet has recorded scores on 21 benchmarks, each shown with its evidence status.
Claude 3.7 Sonnet has recorded scores on 21 benchmarks — Lech Mazur Writing, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, CadEval, Cybench, and 15 more. The full table above shows each score with its evidence status.
5 of 21 recorded scores (24%) are independently reproduced rather than self-reported by the lab.
Claude 3.7 Sonnet has 21 tracked claims: 5 independently reproduced, 16 unverified.
Listed API pricing: $3 per million input tokens, $15 per million output tokens. See the pricing block for the full breakdown.