Anthropic·released 2025-08-051 source
Claude Opus 4.1 benchmark scores: 19 benchmarks tracked. 47% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 19 benchmarks·best result 84.7 on Lech Mazur Writing (unverified)·9 of 19 independently reproduced·$15/$75 per M tokens
Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| Lech Mazur Writing | 84.7accuracy (%) | unverifiedT1 | 2025-08-05lechmazur/writing Github repository |
| SWE-Bench verified | 73.35accuracy (%) | reproduced· optimizedT1 | 2025-08-05Epoch AI |
| GPQA diamond | 69.7accuracy (%) | reproduced· optimizedT1 | 2025-08-05Epoch AI |
| OTIS Mock AIME 2024-2025 | 68.86accuracy (%) | reproduced· optimizedT1 | 2025-08-05Epoch AI |
| METR Time Horizons | 66.81accuracy (%) | unverifiedT1 | 2025-08-05METR - Measuring AI Ability to Complete Long Tasks |
| SimpleBench | 52accuracy (%) | unverifiedT1 | 2025-08-05SimpleBench Leaderboard |
| DeepResearch Bench | 49.7accuracy (%) | unverifiedT1 | 2025-08-05DeepResearchBench Leaderboard |
| GDPval | 43.6accuracy (%) | unverifiedT2 | 2025-08-05 |
| WeirdML | 42.76accuracy (%) | unverifiedT1 | 2025-08-05WeirdML Leaderboard |
| Cybench | 42accuracy (%) | unverified· optimizedT1 | 2025-08-05Cybench leaderboard |
| Terminal Bench | 38accuracy (%) | unverified· optimizedT1 | 2025-08-05Terminal-Bench v2 Leaderboard |
| SimpleQA Verified | 34.8accuracy (%) | reproducedT1 | 2025-08-05Epoch AI |
| Mystery Game Puzzles | 12.98accuracy (%) | reproducedT1 | 2025-08-05Epoch AI |
| FrontierMath-Tiers-1-3-v2-Private | 12.63accuracy (%) | reproducedT1 | 2025-08-05Epoch AI |
| EBR-bench | 7.94accuracy (%) | reproducedT1 | 2025-08-05Epoch AI |
| HLE | 7.06accuracy (%) | unverifiedT2 | 2025-08-05 |
| VPCT | 2.5accuracy (%) | unverifiedT1 | 2025-08-05VPCT leaderboard |
| FrontierMath-Tier-4-v2-Private | 2.44accuracy (%) | reproducedT1 | 2025-08-05Epoch AI |
| Chess Puzzles | 2.15accuracy (%) | reproducedT1 | 2025-08-05Epoch AI |
Claims drawn from cited facts, not live model generation.
Developed by Anthropic, this model was released in 2025.
2 cited facts
This model is scored on 19 tracked benchmarks. It holds no current top score among them. The record is independently reproduced, so these scores are confirmed by outside evaluation.
3 cited facts
With an average gap of 38.52 points behind the leader, this model sits well behind the front of the pack. It holds the top score on zero benchmarks, confirming it has no category leadership.
2 cited facts
6 facts cross-checked across data sources: 3 corroborated, 1 single-source, 2 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, modelsdev_models, openrouter_models
Claude Opus 4.1 is an AI model developed by Anthropic, released 2025-08-05. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Claude Opus 4.1 has recorded scores on 19 benchmarks, each shown with its evidence status.
Claude Opus 4.1 has recorded scores on 19 benchmarks — Lech Mazur Writing, GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, Cybench, and 13 more. The full table above shows each score with its evidence status.
9 of 19 recorded scores (47%) are independently reproduced rather than self-reported by the lab.
Claude Opus 4.1 has 19 tracked claims: 9 independently reproduced, 10 unverified.
Listed API pricing: $15 per million input tokens, $75 per million output tokens. See the pricing block for the full breakdown.