Anthropic·released 2026-02-051 source
Claude Opus 4.6 benchmark scores: 26 benchmarks tracked, leading 3. 35% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 3 of 26 benchmarks·best result 94.44 on OTIS Mock AIME 2024-2025 (reproduced)·9 of 26 independently reproduced·$5/$25 per M tokens
Head-to-headClaude Opus 4.6 vs Claude Opus 4.5Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
Leads on: Cybench, DeepResearch Bench, METR Time Horizons
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 94.44accuracy (%) | reproduced· optimizedT1 | 2026-02-05Epoch AI |
| ARC-AGI | 94accuracy (%) | unverified· optimizedT1 | 2026-02-05https://arcprize.org/leaderboard |
| Cybench | 93accuracy (%) | self-reported· optimizedT1 | 2026-02-05Opus 4.6 System Card |
| GPQA diamond | 87.37accuracy (%) | reproduced· optimizedT1 | 2026-02-05Epoch AI |
| Terminal Bench | 79.8accuracy (%) | unverified· optimizedT1 | 2026-02-05https://www.tbench.ai/leaderboard/terminal-bench/2.0 |
| METR Time Horizons | 78.86accuracy (%) | unverifiedT2 | 2026-02-05 |
| SWE-Bench verified | 78.72accuracy (%) | reproduced· optimizedT1 | 2026-02-05Epoch AI |
| WeirdML | 77.95accuracy (%) | unverifiedT1 | 2026-02-05https://htihle.github.io/weirdml.html |
| ARC-AGI-2 | 69.17accuracy (%) | unverifiedT2 | 2026-02-05 |
| FrontierMath-Tiers-1-3-v2-Private | 65.96accuracy (%) | reproducedT1 | 2026-02-05Epoch AI |
| SimpleBench | 61.12accuracy (%) | unverifiedT1 | 2026-02-05SimpleBench Leaderboard |
| DeepResearch Bench | 55.31accuracy (%) | unverifiedT1 | 2026-02-05https://drb.futuresearch.ai/#drb |
| ProofBench | 50accuracy (%) | unverifiedT1 | 2026-02-05https://www.vals.ai/benchmarks/proof_bench |
| SimpleQA Verified | 46.49accuracy (%) | reproducedT1 | 2026-02-05Epoch AI |
| GSO-Bench | 41.2accuracy (%) | unverifiedT1 | 2026-02-05GSO Leaderboard |
| APEX-Agents | 32.4accuracy (%) | unverifiedT2 | 2026-02-05 |
| HLE | 31.13accuracy (%) | unverifiedT2 | 2026-02-05 |
| FrontierMath-Tier-4-v2-Private | 26.83accuracy (%) | reproducedT1 | 2026-02-05Epoch AI |
| FrontierCode | 26.64accuracy (%) | unverifiedT1 | 2026-02-05https://cognition.com/frontiercode |
| PostTrainBench | 24.82accuracy (%) | unverifiedT2 | 2026-02-05 |
| CL-bench | 20.7accuracy (%) | unverifiedT2 | 2026-02-05 |
| Mystery Game Puzzles | 17.38accuracy (%) | reproducedT1 | 2026-02-05Epoch AI |
| CL-bench Life | 17accuracy (%) | unverifiedT2 | 2026-02-05 |
| EBR-bench | 12.7accuracy (%) | reproducedT1 | 2026-02-05Epoch AI |
| Chess Puzzles | 12.67accuracy (%) | reproducedT1 | 2026-02-05Epoch AI |
| Remote Labor Index | 4.17accuracy (%) | unverifiedT2 | 2026-02-05 |
Claims drawn from cited facts, not live model generation.
This model was developed by Anthropic and released in 2026.
2 cited facts
This model is scored on 26 tracked benchmarks, and it currently tops 3 of them. These results are independently reproduced, so outside evaluation backs the scores up.
3 cited facts
On average across the disclosed benchmark harness, this model trails the SOTA leader by 17.97 points, a wide gap that reads as well behind the front-runners. It nevertheless holds the top score on 3 benchmarks, so on those specific tasks it is a genuine category leader.
2 cited facts
6 facts cross-checked across data sources: 3 corroborated, 2 single-source, 1 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: curated_capability_claims, epoch_benchmarks, litellm_prices, modelsdev_models +1 more
Claude Opus 4.6 is an AI model developed by Anthropic, released 2026-02-05. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Claude Opus 4.6 has recorded scores on 26 benchmarks, each shown with its evidence status.
Claude Opus 4.6 has recorded scores on 26 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, Cybench, SimpleBench, and 20 more. The full table above shows each score with its evidence status.
9 of 26 recorded scores (35%) are independently reproduced rather than self-reported by the lab.
Claude Opus 4.6 has 26 tracked claims: 9 independently reproduced, 1 self-reported, 16 unverified.
Listed API pricing: $5 per million input tokens, $25 per million output tokens. See the pricing block for the full breakdown.