Qwen·released 2026-04-011 source
Qwen 3.6 Plus benchmark scores: 10 benchmarks tracked. 80% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 10 benchmarks·best result 93.33 on OTIS Mock AIME 2024-2025 (reproduced)·8 of 10 independently reproduced·$0.5/$3 per M tokens
Source: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 93.33accuracy (%) | reproduced· optimizedT1 | 2026-04-01Epoch AI |
| GPQA diamond | 84.51accuracy (%) | reproduced· optimizedT1 | 2026-04-01Epoch AI |
| SWE-Bench verified | 57.85accuracy (%) | reproduced· optimizedT1 | 2026-04-01Epoch AI |
| SimpleQA Verified | 49.1accuracy (%) | reproducedT1 | 2026-04-01Epoch AI |
| FrontierMath-2025-02-28-Private | 45.98accuracy (%) | reproduced· optimizedT1 | 2026-04-01Epoch AI |
| CL-bench | 20.3accuracy (%) | unverifiedT2 | 2026-04-01 |
| Mystery Game Puzzles | 19.59accuracy (%) | reproducedT1 | 2026-04-01Epoch AI |
| FrontierMath-Tier-4-2025-07-01-Private | 13.89accuracy (%) | reproducedT1 | 2026-04-01Epoch AI |
| Chess Puzzles | 12.67accuracy (%) | reproducedT1 | 2026-04-01Epoch AI |
| CritPt | 2.86accuracy (%) | unverifiedT2 | 2026-04-01 |
Claims drawn from cited facts, not live model generation.
This model was developed by Qwen and released in 2026.
2 cited facts
This model is scored on 10 tracked benchmarks, yet it holds no current top score on any of them. The record is independently reproduced, so the scores are backed by outside evaluation and can be treated as verified.
3 cited facts
With an average gap of 23.06 points behind the state of the art, this model sits well behind the leaders rather than at the frontier. It holds the top score on none of the benchmark comparisons (0 leader positions), confirming it has no category leadership yet.
2 cited facts
5 facts cross-checked across data sources: 2 corroborated, 1 single-source, 2 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, modelsdev_models, openrouter_models
Qwen 3.6 Plus is an AI model developed by Qwen, released 2026-04-01. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Qwen 3.6 Plus has recorded scores on 10 benchmarks, each shown with its evidence status.
Qwen 3.6 Plus has recorded scores on 10 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, Chess Puzzles, FrontierMath-2025-02-28-Private, SimpleQA Verified, CritPt, and 4 more. The full table above shows each score with its evidence status.
8 of 10 recorded scores (80%) are independently reproduced rather than self-reported by the lab.
Qwen 3.6 Plus has 10 tracked claims: 8 independently reproduced, 2 unverified.
Listed API pricing: $0.5 per million input tokens, $3 per million output tokens (prices disputed across sources). See the pricing block for the full breakdown.