Anthropic·released 2026-02-171 source
Claude Sonnet 4.6 benchmark scores: 22 benchmarks tracked, leading 1. 36% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 1 of 22 benchmarks·best result 86.5 on ARC-AGI (unverified)·8 of 22 independently reproduced·$3/$15 per M tokens
Head-to-headClaude Sonnet 4.6 vs Claude Opus 4.5Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
Leads on: OSWorld
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| ARC-AGI | 86.5accuracy (%) | unverified· optimizedT2 | 2026-02-17 |
| OTIS Mock AIME 2024-2025 | 85.78accuracy (%) | reproduced· optimizedT1 | 2026-02-17Epoch AI |
| GPQA diamond | 83.16accuracy (%) | reproduced· optimizedT1 | 2026-02-17Epoch AI |
| SWE-Bench verified | 75.21accuracy (%) | reproduced· optimizedT1 | 2026-02-17Epoch AI |
| OSWorld | 72.1accuracy (%) | unverified· optimizedT1 | 2026-02-17OS World Website |
| WeirdML | 66.07accuracy (%) | unverifiedT1 | 2026-02-17WeirdML Leaderboard |
| ARC-AGI-2 | 60.42accuracy (%) | unverifiedT2 | 2026-02-17 |
| FrontierMath-2025-02-28-Private | 56.84accuracy (%) | reproduced· optimizedT1 | 2026-02-17Epoch AI |
| DeepResearch Bench | 54.87accuracy (%) | unverifiedT1 | 2026-02-17https://drb.futuresearch.ai/#drb |
| Terminal Bench | 53.4accuracy (%) | unverified· optimizedT1 | 2026-02-17https://www.tbench.ai/leaderboard/terminal-bench/2.0 |
| ProofBench | 45accuracy (%) | unverifiedT1 | 2026-02-17https://www.vals.ai/benchmarks/proof_bench |
| DeepSWE | 29.93accuracy (%) | unverifiedT1 | 2026-02-17https://deepswe.datacurve.ai/ |
| SimpleQA Verified | 29accuracy (%) | reproducedT1 | 2026-02-17Epoch AI |
| FrontierCode | 24.31accuracy (%) | unverifiedT1 | 2026-02-17https://cognition.com/frontiercode |
| APEX-Agents | 23.7accuracy (%) | unverifiedT2 | 2026-02-17 |
| ExploitBench | 23.6accuracy (%) | unverifiedT2 | 2026-02-17 |
| PostTrainBench | 16.42accuracy (%) | unverifiedT2 | 2026-02-17 |
| FrontierMath-Tier-4-2025-07-01-Private | 13.83accuracy (%) | reproducedT1 | 2026-02-17Epoch AI |
| OSWorld 2.0 | 9.3accuracy (%) | unverifiedT1 | 2026-02-17https://osworld-v2.xlang.ai/ |
| Chess Puzzles | 8.46accuracy (%) | reproducedT1 | 2026-02-17Epoch AI |
| Mystery Game Puzzles | 7.47accuracy (%) | reproducedT1 | 2026-02-17Epoch AI |
| CritPt | 3.14accuracy (%) | unverifiedT2 | 2026-02-17 |
Claims drawn from cited facts, not live model generation.
Developed by Anthropic, this model was released in 2026.
2 cited facts
This model is scored on 22 tracked benchmarks, and it currently holds the top score on 1 of them. Because the scores have been independently reproduced by outside evaluation, the record is trustworthy.
3 cited facts
With an average gap of 25.01 points to the best scores, this model sits well behind the leaders on its benchmarks. It holds the top score on 1 benchmark, but that one leadership spot does not offset the wide overall gap.
3 cited facts
6 facts cross-checked across data sources: 3 corroborated, 2 single-source, 1 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, modelsdev_models, openrouter_models
Claude Sonnet 4.6 is an AI model developed by Anthropic, released 2026-02-17. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Claude Sonnet 4.6 has recorded scores on 22 benchmarks, each shown with its evidence status.
Claude Sonnet 4.6 has recorded scores on 22 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, FrontierMath-2025-02-28-Private, SimpleQA Verified, and 16 more. The full table above shows each score with its evidence status.
8 of 22 recorded scores (36%) are independently reproduced rather than self-reported by the lab.
Claude Sonnet 4.6 has 22 tracked claims: 8 independently reproduced, 14 unverified.
Listed API pricing: $3 per million input tokens, $15 per million output tokens. See the pricing block for the full breakdown.