Qwen·released 2026-07-191 source
Qwen 3.8 Max benchmark scores: 9 benchmarks tracked. 78% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 9 benchmarks·best result 99.44 on OTIS Mock AIME 2024-2025 (reproduced)·7 of 9 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 99.44accuracy (%) | reproduced· optimizedT1 | 2026-07-19Epoch AI |
| GPQA diamond | 90.24accuracy (%) | reproduced· optimizedT1 | 2026-07-19Epoch AI |
| FrontierMath-Tiers-1-3-v2-Private | 74.74accuracy (%) | reproducedT1 | 2026-07-19Epoch AI |
| ProofBench | 58accuracy (%) | unverifiedT1 | 2026-07-19https://www.vals.ai/benchmarks/proof_bench |
| DeepSWE | 57.46accuracy (%) | unverifiedT1 | 2026-07-19https://deepswe.datacurve.ai/ |
| FrontierMath-Tier-4-v2-Private | 46.34accuracy (%) | reproducedT1 | 2026-07-19Epoch AI |
| SimpleQA Verified | 46.29accuracy (%) | reproducedT1 | 2026-07-19Epoch AI |
| Mystery Game Puzzles | 31.7accuracy (%) | reproducedT1 | 2026-07-19Epoch AI |
| Chess Puzzles | 25.29accuracy (%) | reproducedT1 | 2026-07-19Epoch AI |
Claims drawn from cited facts, not live model generation.
Developed by Qwen and released in 2026, this model's parameter scale is not disclosed.
2 cited facts
This model is evaluated on nine tracked benchmarks and currently holds no top score in any of them. The record is independently reproduced, so outside evaluation corroborates these results.
3 cited facts
It trails the leading score on its benchmarks by an average of 23.05 points, a wide gap that places it well behind the leaders. It holds the top score on none of the tracked benchmarks, so it has no category leadership yet.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Qwen 3.8 Max is an AI model developed by Qwen, released 2026-07-19. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Qwen 3.8 Max has recorded scores on 9 benchmarks, each shown with its evidence status.
Qwen 3.8 Max has recorded scores on 9 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, Chess Puzzles, SimpleQA Verified, FrontierMath-Tiers-1-3-v2-Private, FrontierMath-Tier-4-v2-Private, and 3 more. The full table above shows each score with its evidence status.
7 of 9 recorded scores (78%) are independently reproduced rather than self-reported by the lab.
Qwen 3.8 Max has 9 tracked claims: 7 independently reproduced, 2 unverified.