Qwen·released 2026-09-011 source
Qwen3.8 Max (0902) benchmark scores: 6 benchmarks tracked, leading 1. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 1 of 6 benchmarks·best result 100 on OTIS Mock AIME 2024-2025 (unverified)·0 of 6 independently reproduced
Head-to-headQwen3.8 Max (0902) vs GPT-5.5Leads on: OTIS Mock AIME 2024-2025
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 100accuracy (%) | unverifiedT2 | 2026-09-01 |
| GPQA diamond | 89.73accuracy (%) | unverifiedT2 | 2026-09-01 |
| FrontierMath-Tiers-1-3-v2-Private | 65.61accuracy (%) | unverifiedT2 | 2026-09-01 |
| SimpleQA Verified | 47.3accuracy (%) | unverifiedT2 | 2026-09-01 |
| Chess Puzzles | 36.87accuracy (%) | unverifiedT2 | 2026-09-01 |
| FrontierMath-Tier-4-v2-Private | 34.15accuracy (%) | unverifiedT2 | 2026-09-01 |
Claims drawn from cited facts, not live model generation.
This model was developed by Qwen, placing its provenance with an established large-scale model builder. It arrives with a 2026 release vintage, making it a recent entry whose recency reflects the current generation of tooling and training practice rather than an older lineage.
2 cited facts
It is scored on 6 tracked benchmarks and currently holds the top score on 1 of them. Its record is unverified: outside evaluation has neither confirmed nor disputed these numbers.
3 cited facts
Across its benchmarks the model trails the leading score by 26.35 points on average, a wide gap that reads as well behind the front of the pack rather than at the frontier or merely competitive. It holds the top score on only 1 benchmark, so its best result amounts to a single isolated win rather than genuine category leadership.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Qwen3.8 Max (0902) is an AI model developed by Qwen, released 2026-09-01. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Qwen3.8 Max (0902) has recorded scores on 6 benchmarks, each shown with its evidence status.
Qwen3.8 Max (0902) has recorded scores on 6 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, SimpleQA Verified, Chess Puzzles, FrontierMath-Tiers-1-3-v2-Private, FrontierMath-Tier-4-v2-Private. The full table above shows each score with its evidence status.
0 of 6 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Qwen3.8 Max (0902) has 6 tracked claims: 6 unverified.