Qwen·released 2026-02-161 source
Qwen 3.5 Plus (hosted 397B-A17B) benchmark scores: 10 benchmarks tracked. 70% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 10 benchmarks·best result 86.65 on OTIS Mock AIME 2024-2025 (reproduced)·7 of 10 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 86.65accuracy (%) | reproduced· optimizedT1 | 2026-02-16Epoch AI |
| GPQA diamond | 79.8accuracy (%) | reproduced· optimizedT1 | 2026-02-16Epoch AI |
| FrontierMath-2025-02-28-Private | 36.9accuracy (%) | reproduced· optimizedT1 | 2026-02-16Epoch AI |
| SimpleQA Verified | 26accuracy (%) | reproducedT1 | 2026-02-16Epoch AI |
| CL-bench | 19.8accuracy (%) | unverifiedT2 | 2026-02-16 |
| Chess Puzzles | 17.93accuracy (%) | reproducedT1 | 2026-02-16Epoch AI |
| APEX-Agents | 13.6accuracy (%) | unverifiedT2 | 2026-02-16 |
| CL-bench Life | 12.4accuracy (%) | unverifiedT2 | 2026-02-16 |
| Mystery Game Puzzles | 7.47accuracy (%) | reproducedT1 | 2026-02-16Epoch AI |
| FrontierMath-Tier-4-2025-07-01-Private | 3.47accuracy (%) | reproducedT1 | 2026-02-16Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by Qwen and released in 2026.
2 cited facts
This model is scored on 10 tracked benchmarks and holds no current top score. Its record is independently reproduced, so outside evaluation backs the reported scores up.
3 cited facts
Based on the average gap of 27.81 points, this model trails the leader by a wide margin, placing it well behind the front-of-pack on benchmark scores. It holds the top score on none of the tracked benchmarks, so it shows no current category leadership.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Qwen 3.5 Plus (hosted 397B-A17B) is an AI model developed by Qwen, released 2026-02-16. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Qwen 3.5 Plus (hosted 397B-A17B) has recorded scores on 10 benchmarks, each shown with its evidence status.
Qwen 3.5 Plus (hosted 397B-A17B) has recorded scores on 10 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, Chess Puzzles, FrontierMath-2025-02-28-Private, SimpleQA Verified, FrontierMath-Tier-4-2025-07-01-Private, and 4 more. The full table above shows each score with its evidence status.
7 of 10 recorded scores (70%) are independently reproduced rather than self-reported by the lab.
Qwen 3.5 Plus (hosted 397B-A17B) has 10 tracked claims: 7 independently reproduced, 3 unverified.