Qwen·released 2026-02-131 source
Qwen3.5 397B-A17B benchmark scores: 7 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 7 benchmarks·best result 88.88 on OTIS Mock AIME 2024-2025 (unverified)·0 of 7 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 88.88accuracy (%) | unverifiedT2 | 2026-02-13 |
| GPQA diamond | 81.82accuracy (%) | unverifiedT2 | 2026-02-13 |
| DTBench | 79.12accuracy (%) | unverifiedT1 | 2026-02-13https://conceptualreasoning.ai/dtbench |
| LMCA | 44.58accuracy (%) | unverifiedT1 | 2026-02-13https://conceptualreasoning.ai/lmca |
| FrontierMath-Tiers-1-3-v2-Private | 31.23accuracy (%) | unverifiedT2 | 2026-02-13 |
| Mystery Game Puzzles | 9.67accuracy (%) | unverifiedT2 | 2026-02-13 |
| Chess Puzzles | 8.46accuracy (%) | unverifiedT2 | 2026-02-13 |
Claims drawn from cited facts, not live model generation.
This model comes from Qwen, placing its origin among the developer's own model line. It reflects a recent vintage, having been released in 2026, which situates it within the developer's newer generation of work.
3 cited facts
This model is scored on seven tracked benchmarks, and it currently holds no top score on any of them. Its record remains unverified: neither confirmed nor disputed by outside evaluation, so the reported figures should be read as provisional claims rather than settled results.
3 cited facts
On average this model trails the SOTA leader by 38.43 points across its benchmarks, a wide gap that reads as well behind the front of the pack rather than at or near the frontier — a score gap under whatever harness is disclosed, not a statement about capability. Consistently, it holds the top score on 0 benchmarks, meaning there is no benchmark here where it is the category leader.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Qwen3.5 397B-A17B is an AI model developed by Qwen, released 2026-02-13. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Qwen3.5 397B-A17B has recorded scores on 7 benchmarks, each shown with its evidence status.
Qwen3.5 397B-A17B has recorded scores on 7 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, DTBench, LMCA, Chess Puzzles, FrontierMath-Tiers-1-3-v2-Private, and 1 more. The full table above shows each score with its evidence status.
0 of 7 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Qwen3.5 397B-A17B has 7 tracked claims: 7 unverified.