Qwen·released 2026-06-021 source
Qwen3.7-Plus benchmark scores: 6 benchmarks tracked. 50% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 6 benchmarks·best result 93.33 on OTIS Mock AIME 2024-2025 (reproduced)·3 of 6 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 93.33accuracy (%) | reproduced· optimizedT1 | 2026-06-02Epoch AI |
| GPQA diamond | 83.84accuracy (%) | reproduced· optimizedT1 | 2026-06-02Epoch AI |
| Chess Puzzles | 20.03accuracy (%) | reproducedT1 | 2026-06-02Epoch AI |
| FrontierCode | 10.2accuracy (%) | unverifiedT1 | 2026-06-02https://cognition.com/frontiercode |
| CritPt | 9.14accuracy (%) | unverifiedT2 | 2026-06-02 |
| OSWorld 2.0 | 2.8accuracy (%) | unverifiedT1 | 2026-06-02https://osworld-v2.xlang.ai/ |
Claims drawn from cited facts, not live model generation.
This model was developed by Qwen and released in 2026.
2 cited facts
Across tracked benchmarks, this model is scored on 6, yet it holds no current top score in any of them. Because these results are independently reproduced by outside evaluation, the record can be trusted as verified rather than vendor-claimed.
3 cited facts
On average it trails the SOTA leader by 23.71 points under the disclosed harness, a wide gap that places it well behind the leaders. It holds the top score on none of the tracked benchmarks, so it has no category leadership yet.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Qwen3.7-Plus is an AI model developed by Qwen, released 2026-06-02. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Qwen3.7-Plus has recorded scores on 6 benchmarks, each shown with its evidence status.
Qwen3.7-Plus has recorded scores on 6 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, Chess Puzzles, CritPt, FrontierCode, OSWorld 2.0. The full table above shows each score with its evidence status.
3 of 6 recorded scores (50%) are independently reproduced rather than self-reported by the lab.
Qwen3.7-Plus has 6 tracked claims: 3 independently reproduced, 3 unverified.