Qwen·released 2026-04-141 source
Qwen 3.6 35B-A3B benchmark scores: 7 benchmarks tracked. 43% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 7 benchmarks·best result 86.65 on OTIS Mock AIME 2024-2025 (reproduced)·3 of 7 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 86.65accuracy (%) | reproduced· optimizedT1 | 2026-04-14Epoch AI |
| GPQA diamond | 79.8accuracy (%) | reproduced· optimizedT1 | 2026-04-14Epoch AI |
| Surface Evolver Bench | 44.38accuracy (%) | unverifiedT1 | 2026-04-14https://yhenon.github.io/surface-evolver-llm-eval/ |
| WeirdML | 34.49accuracy (%) | unverifiedT1 | 2026-04-14https://htihle.github.io/weirdml.html |
| Terminal Bench | 24.6accuracy (%) | unverified· optimizedT1 | 2026-04-14https://www.tbench.ai/leaderboard/terminal-bench/2.0 |
| Chess Puzzles | 22.14accuracy (%) | reproducedT1 | 2026-04-14Epoch AI |
| CritPt | 0.29accuracy (%) | unverifiedT2 | 2026-04-14 |
Claims drawn from cited facts, not live model generation.
Developed by Qwen, this model was released in 2026.
2 cited facts
This model is scored on seven tracked benchmarks, but it currently holds no top score in any of them. Since its scores have been independently reproduced, the record can be treated as reliable.
3 cited facts
On average it trails the benchmark leader by 38.12 points, a wide gap that places it well behind the front of the pack. It holds the top score on none of the tracked leaderboards, so no genuine category leadership is established.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Qwen 3.6 35B-A3B is an AI model developed by Qwen, released 2026-04-14. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Qwen 3.6 35B-A3B has recorded scores on 7 benchmarks, each shown with its evidence status.
Qwen 3.6 35B-A3B has recorded scores on 7 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, CritPt, Surface Evolver Bench, and 1 more. The full table above shows each score with its evidence status.
3 of 7 recorded scores (43%) are independently reproduced rather than self-reported by the lab.
Qwen 3.6 35B-A3B has 7 tracked claims: 3 independently reproduced, 4 unverified.