Qwen·released 2026-05-191 source
Qwen3.7-Max benchmark scores: 12 benchmarks tracked. 75% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 12 benchmarks·best result 95.55 on OTIS Mock AIME 2024-2025 (reproduced)·9 of 12 independently reproduced·$2.5/$7.5 per M tokens
Source: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 95.55accuracy (%) | reproduced· optimizedT1 | 2026-05-19Epoch AI |
| GPQA diamond | 87.88accuracy (%) | reproduced· optimizedT1 | 2026-05-19Epoch AI |
| SWE-Bench verified | 77.27accuracy (%) | reproduced· optimizedT1 | 2026-05-19Epoch AI |
| FrontierMath-Tiers-1-3-v2-Private | 64.56accuracy (%) | reproducedT1 | 2026-05-19Epoch AI |
| SimpleBench | 64.48accuracy (%) | unverifiedT1 | 2026-05-19SimpleBench Leaderboard |
| SimpleQA Verified | 58.52accuracy (%) | reproducedT1 | 2026-05-19Epoch AI |
| FrontierMath-Tier-4-v2-Private | 34.15accuracy (%) | reproducedT1 | 2026-05-19Epoch AI |
| ProofBench | 26accuracy (%) | unverifiedT1 | 2026-05-19https://www.vals.ai/benchmarks/proof_bench |
| Mystery Game Puzzles | 25.09accuracy (%) | reproducedT1 | 2026-05-19Epoch AI |
| Chess Puzzles | 14.77accuracy (%) | reproducedT1 | 2026-05-19Epoch AI |
| CritPt | 13.43accuracy (%) | unverifiedT2 | 2026-05-19 |
| EBR-bench | 9.52accuracy (%) | reproducedT1 | 2026-05-19Epoch AI |
Claims drawn from cited facts, not live model generation.
This model originates from Qwen and was released in 2026.
2 cited facts
This model is scored on 12 tracked benchmarks and currently holds no top score in any of them. Because these results have been independently reproduced, readers can treat the benchmark record as verified rather than as vendor-only claims.
3 cited facts
Relative to the disclosed benchmark harness, this model trails the average SOTA score by 28.01 points, a wide gap that places it well behind the leaders. It holds the top score on none of the tracked benchmarks, so there is no genuine category leadership to report.
2 cited facts
5 facts cross-checked across data sources: 2 corroborated, 1 single-source, 2 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, modelsdev_models, openrouter_models
Qwen3.7-Max is an AI model developed by Qwen, released 2026-05-19. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Qwen3.7-Max has recorded scores on 12 benchmarks, each shown with its evidence status.
Qwen3.7-Max has recorded scores on 12 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, Chess Puzzles, SimpleBench, SimpleQA Verified, CritPt, and 6 more. The full table above shows each score with its evidence status.
9 of 12 recorded scores (75%) are independently reproduced rather than self-reported by the lab.
Qwen3.7-Max has 12 tracked claims: 9 independently reproduced, 3 unverified.
Listed API pricing: $2.5 per million input tokens, $7.5 per million output tokens (prices disputed across sources). See the pricing block for the full breakdown.