Qwen·released 2026-04-201 source
Qwen 3.6 Max (Preview) benchmark scores: 9 benchmarks tracked. 89% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 9 benchmarks·best result 91.1 on OTIS Mock AIME 2024-2025 (reproduced)·8 of 9 independently reproduced·$1.3/$7.8 per M tokens
Source: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 91.1accuracy (%) | reproduced· optimizedT1 | 2026-04-20Epoch AI |
| GPQA diamond | 83.16accuracy (%) | reproduced· optimizedT1 | 2026-04-20Epoch AI |
| SWE-Bench verified | 76.65accuracy (%) | reproduced· optimizedT1 | 2026-04-20Epoch AI |
| SimpleQA Verified | 56.93accuracy (%) | reproducedT1 | 2026-04-20Epoch AI |
| SimpleBench | 55.6accuracy (%) | unverifiedT1 | 2026-04-20SimpleBench Leaderboard |
| FrontierMath-2025-02-28-Private | 40.53accuracy (%) | reproduced· optimizedT1 | 2026-04-20Epoch AI |
| Chess Puzzles | 15.82accuracy (%) | reproducedT1 | 2026-04-20Epoch AI |
| Mystery Game Puzzles | 7.47accuracy (%) | reproducedT1 | 2026-04-20Epoch AI |
| FrontierMath-Tier-4-2025-07-01-Private | 6.94accuracy (%) | reproducedT1 | 2026-04-20Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by Qwen and released in 2026.
2 cited facts
The model is scored on 9 tracked benchmarks. It holds no current top score among those benchmarks. Because these scores have been independently reproduced by outside evaluation, the record can be trusted as verified.
3 cited facts
With an average gap of 23.84 points to the state of the art, this model sits well behind the leaders. It holds the top score on none of the tracked benchmarks, so it has no category leadership yet.
2 cited facts
5 facts cross-checked across data sources: 2 corroborated, 1 single-source, 2 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: curated_capability_claims, epoch_benchmarks, modelsdev_models, openrouter_models
Qwen 3.6 Max (Preview) is an AI model developed by Qwen, released 2026-04-20. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Qwen 3.6 Max (Preview) has recorded scores on 9 benchmarks, each shown with its evidence status.
Qwen 3.6 Max (Preview) has recorded scores on 9 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, Chess Puzzles, SimpleBench, FrontierMath-2025-02-28-Private, SimpleQA Verified, and 3 more. The full table above shows each score with its evidence status.
8 of 9 recorded scores (89%) are independently reproduced rather than self-reported by the lab.
Qwen 3.6 Max (Preview) has 9 tracked claims: 8 independently reproduced, 1 unverified.
Listed API pricing: $1.3 per million input tokens, $7.8 per million output tokens (prices disputed across sources). See the pricing block for the full breakdown.