Anthropic·released 2026-05-281 source
Claude Opus 4.8 benchmark scores: 23 benchmarks tracked, leading 2. 35% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 2 of 23 benchmarks·best result 98.33 on OTIS Mock AIME 2024-2025 (reproduced)·8 of 23 independently reproduced·$5/$25 per M tokens
Head-to-headClaude Opus 4.8 vs Claude Opus 4.7Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
Leads on: GSO-Bench, OSWorld 2.0
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 98.33accuracy (%) | reproduced· optimizedT1 | 2026-05-28Epoch AI |
| ARC-AGI | 92.5accuracy (%) | unverified· optimizedT1 | 2026-05-28https://arcprize.org/leaderboard |
| GPQA diamond | 88.05accuracy (%) | reproduced· optimizedT1 | 2026-05-28Epoch AI |
| Surface Evolver Bench | 87.5accuracy (%) | unverifiedT1 | 2026-05-28https://yhenon.github.io/surface-evolver-llm-eval/ |
| WeirdML | 82.89accuracy (%) | unverifiedT1 | 2026-05-28https://htihle.github.io/weirdml.html |
| FrontierMath-Tiers-1-3-v2-Private | 80accuracy (%) | reproducedT1 | 2026-05-28Epoch AI |
| ARC-AGI-2 | 72.08accuracy (%) | unverifiedT2 | 2026-05-28 |
| ProofBench | 69accuracy (%) | unverifiedT1 | 2026-05-28https://www.vals.ai/benchmarks/proof_bench |
| DeepSWE | 58.97accuracy (%) | unverifiedT1 | 2026-05-28https://deepswe.datacurve.ai/ |
| SimpleBench | 57.76accuracy (%) | unverifiedT1 | 2026-05-28https://lmcouncil.ai/benchmarks |
| FrontierMath-Tier-4-v2-Private | 56.1accuracy (%) | reproducedT1 | 2026-05-28Epoch AI |
| DeepResearch Bench | 50.23accuracy (%) | unverifiedT1 | 2026-05-28https://drb.futuresearch.ai/#drb |
| GSO-Bench | 47.06accuracy (%) | unverifiedT1 | 2026-05-28https://gso-bench.github.io/index.html |
| FrontierCode | 46.5accuracy (%) | unverifiedT1 | 2026-05-28https://cognition.com/frontiercode |
| APEX-Agents | 42.5accuracy (%) | unverifiedT2 | 2026-05-28 |
| SimpleQA Verified | 39.5accuracy (%) | reproducedT1 | 2026-05-28Epoch AI |
| PostTrainBench | 34.08accuracy (%) | unverifiedT2 | 2026-05-28 |
| Chess Puzzles | 30.56accuracy (%) | reproducedT1 | 2026-05-28Epoch AI |
| Mystery Game Puzzles | 29.5accuracy (%) | reproducedT1 | 2026-05-28Epoch AI |
| EBR-bench | 28.57accuracy (%) | reproducedT1 | 2026-05-28Epoch AI |
| CritPt | 20.86accuracy (%) | unverifiedT2 | 2026-05-28 |
| OSWorld 2.0 | 20.6accuracy (%) | unverifiedT1 | 2026-05-28https://osworld-v2.xlang.ai/ |
| Remote Labor Index | 8.33accuracy (%) | unverifiedT2 | 2026-05-28 |
Claims drawn from cited facts, not live model generation.
This model was developed by Anthropic and released in 2026.
2 cited facts
This model is scored on 23 tracked benchmarks and currently holds the top score on 2 of them. Because these results have been independently reproduced by outside evaluation, the record should be treated as verified rather than merely vendor-claimed.
3 cited facts
This model trails the top score by an average of 13.62 points on the disclosed harness, a wide margin that places it well behind the leaders rather than at the frontier. Despite that average gap, it still holds the leading score on 2 benchmarks, indicating isolated strengths rather than overall category leadership.
2 cited facts
6 facts cross-checked across data sources: 3 corroborated, 2 single-source, 1 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, modelsdev_models, openrouter_models
Claude Opus 4.8 is an AI model developed by Anthropic, released 2026-05-28. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Claude Opus 4.8 has recorded scores on 23 benchmarks, each shown with its evidence status.
Claude Opus 4.8 has recorded scores on 23 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleBench, SimpleQA Verified, and 17 more. The full table above shows each score with its evidence status.
8 of 23 recorded scores (35%) are independently reproduced rather than self-reported by the lab.
Claude Opus 4.8 has 23 tracked claims: 8 independently reproduced, 15 unverified.
Listed API pricing: $5 per million input tokens, $25 per million output tokens. See the pricing block for the full breakdown.