Qwen·released 2025-09-051 source
Qwen3-Max benchmark scores: 8 benchmarks tracked. 63% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 8 benchmarks·best result 97.13 on MATH level 5 (reproduced)·5 of 8 independently reproduced·$2.11/$8.45 per M tokens
Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MATH level 5 | 97.13accuracy (%) | reproduced· optimizedT1 | 2025-09-05Epoch AI |
| OTIS Mock AIME 2024-2025 | 73.31accuracy (%) | reproduced· optimizedT1 | 2025-09-05Epoch AI |
| SimpleQA Verified | 67.47accuracy (%) | reproducedT1 | 2025-09-05Epoch AI |
| Fiction.LiveBench | 66.7accuracy (%) | unverifiedT1 | 2025-09-05Fiction.live leaderboard |
| GPQA diamond | 63.47accuracy (%) | reproduced· optimizedT1 | 2025-09-05Epoch AI |
| CL-bench | 14.5accuracy (%) | unverifiedT2 | 2025-09-05 |
| PostTrainBench | 7.42accuracy (%) | unverifiedT2 | 2025-09-05 |
| Chess Puzzles | 0accuracy (%) | reproducedT1 | 2025-09-05Epoch AI |
Claims drawn from cited facts, not live model generation.
This model is tracked across 8 benchmarks and currently holds no top score on any of them. Because the record has been independently reproduced, outside evaluation supports these results.
3 cited facts
The model trails the average state-of-the-art score by 25.94 points under the disclosed harness, a wide gap that places it well behind the leaders. It holds the top score on none of the benchmark comparisons, so it has no genuine category leadership yet.
2 cited facts
5 facts cross-checked across data sources: 1 single-source, 4 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, modelsdev_models, openrouter_models
Qwen3-Max is an AI model developed by Qwen, released 2025-09-05. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Qwen3-Max has recorded scores on 8 benchmarks, each shown with its evidence status.
Qwen3-Max has recorded scores on 8 benchmarks — GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, Chess Puzzles, SimpleQA Verified, Fiction.LiveBench, and 2 more. The full table above shows each score with its evidence status.
5 of 8 recorded scores (63%) are independently reproduced rather than self-reported by the lab.
Qwen3-Max has 8 tracked claims: 5 independently reproduced, 3 unverified.
Listed API pricing: $2.11 per million input tokens, $8.45 per million output tokens (prices disputed across sources). See the pricing block for the full breakdown.