Qwen·released 2026-04-261 source
Qwen 3.6 Flash benchmark scores: 8 benchmarks tracked. 88% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 8 benchmarks·best result 84.43 on OTIS Mock AIME 2024-2025 (reproduced)·7 of 8 independently reproduced·$0.188/$1.13 per M tokens
Source: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 84.43accuracy (%) | reproduced· optimizedT1 | 2026-04-26Epoch AI |
| GPQA diamond | 77.78accuracy (%) | reproduced· optimizedT1 | 2026-04-26Epoch AI |
| SimpleBench | 22.24accuracy (%) | unverifiedT1 | 2026-04-26SimpleBench Leaderboard |
| SimpleQA Verified | 21.16accuracy (%) | reproducedT1 | 2026-04-26Epoch AI |
| FrontierMath-2025-02-28-Private | 18.15accuracy (%) | reproduced· optimizedT1 | 2026-04-26Epoch AI |
| Chess Puzzles | 15.82accuracy (%) | reproducedT1 | 2026-04-26Epoch AI |
| Mystery Game Puzzles | 5.27accuracy (%) | reproducedT1 | 2026-04-26Epoch AI |
| FrontierMath-Tier-4-2025-07-01-Private | 0accuracy (%) | reproducedT1 | 2026-04-26Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by Qwen and released in 2026.
2 cited facts
This model is scored on 8 tracked benchmarks, and it currently holds no top score on any of them. Because the record is independently reproduced, outside evaluation backs these results, so the scores can be treated as verified rather than vendor claims.
3 cited facts
With an average gap of 40.06 points to the state of the art across benchmarks, this model trails the leaders by a wide margin rather than being competitive at the frontier. It holds the top score on none of the tracked benchmarks, so there is no genuine category leadership yet.
2 cited facts
5 facts cross-checked across data sources: 4 corroborated, 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, modelsdev_models, openrouter_models
Qwen 3.6 Flash is an AI model developed by Qwen, released 2026-04-26. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Qwen 3.6 Flash has recorded scores on 8 benchmarks, each shown with its evidence status.
Qwen 3.6 Flash has recorded scores on 8 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, Chess Puzzles, SimpleBench, FrontierMath-2025-02-28-Private, SimpleQA Verified, and 2 more. The full table above shows each score with its evidence status.
7 of 8 recorded scores (88%) are independently reproduced rather than self-reported by the lab.
Qwen 3.6 Flash has 8 tracked claims: 7 independently reproduced, 1 unverified.
Listed API pricing: $0.1875 per million input tokens, $1.125 per million output tokens. See the pricing block for the full breakdown.