Qwen·released 2026-02-241 source
Qwen3.5-9B benchmark scores: 6 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 6 benchmarks·best result 71.97 on GPQA diamond (unverified)·0 of 6 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| GPQA diamond | 71.97accuracy (%) | unverifiedT2 | 2026-02-24 |
| OTIS Mock AIME 2024-2025 | 61.63accuracy (%) | unverifiedT2 | 2026-02-24 |
| DTBench | 52accuracy (%) | unverifiedT1 | 2026-02-24https://conceptualreasoning.ai/dtbench |
| LMCA | 28.79accuracy (%) | unverifiedT1 | 2026-02-24https://conceptualreasoning.ai/lmca |
| Terminal Bench | 9.2accuracy (%) | unverifiedT1 | 2026-02-24https://www.tbench.ai/leaderboard/terminal-bench/2.0 |
| Chess Puzzles | 7.41accuracy (%) | unverifiedT2 | 2026-02-24 |
Claims drawn from cited facts, not live model generation.
This model was developed by Qwen and released in 2026, placing it in a recent vintage of models from that team.
2 cited facts
This model is scored on 6 tracked benchmarks, and it holds no current top score on any of them. That record is unverified — neither confirmed nor disputed by independent outside evaluation.
3 cited facts
On average this model trails the best reported score by 48.4 points under a disclosed harness — a wide gap that reads as well behind the leaders rather than a narrow near-frontier miss, and it is a score differential, not a capability verdict. It holds the top score on none of its benchmarks, so there is no category in which it currently leads the field.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Qwen3.5-9B is an AI model developed by Qwen, released 2026-02-24. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Qwen3.5-9B has recorded scores on 6 benchmarks, each shown with its evidence status.
Qwen3.5-9B has recorded scores on 6 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, DTBench, LMCA, Chess Puzzles, Terminal Bench. The full table above shows each score with its evidence status.
0 of 6 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Qwen3.5-9B has 6 tracked claims: 6 unverified.