Qwen·released 2025-01-251 source
Qwen2.5-Max benchmark scores: 7 benchmarks tracked. 57% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 7 benchmarks·best result 72.9 on Lech Mazur Writing (unverified)·4 of 7 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| Lech Mazur Writing | 72.9accuracy (%) | unverifiedT1 | 2025-01-25lechmazur/writing Github repository |
| MATH level 5 | 67.18accuracy (%) | reproduced· optimizedT1 | 2025-01-25Epoch AI |
| Fiction.LiveBench | 66.7accuracy (%) | unverifiedT1 | 2025-01-25Fiction.live leaderboard |
| GPQA diamond | 41.5accuracy (%) | reproduced· optimizedT1 | 2025-01-25Epoch AI |
| Aider polyglot | 21.8accuracy (%) | unverified· optimizedT1 | 2025-01-25Aider LLM Leaderboards |
| OTIS Mock AIME 2024-2025 | 16.03accuracy (%) | reproduced· optimizedT1 | 2025-01-25Epoch AI |
| FrontierMath-2025-02-28-Private | 1.81accuracy (%) | reproduced· optimizedT1 | 2025-01-25Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by Qwen. It was released in 2025.
2 cited facts
Off the front across all 7 of its measured benchmarks, this model earns a field placement — valid for these harnesses, silent beyond them. Because the scores are independently reproduced, the numbers are reliable and have been confirmed by outside evaluation.
3 cited facts
On average, this model trails the state-of-the-art by 49.13 points, placing it in the 'wide gap' band—well behind the leading models.
1 cited fact
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Qwen2.5-Max is an AI model developed by Qwen, released 2025-01-25. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Qwen2.5-Max has recorded scores on 7 benchmarks, each shown with its evidence status.
Qwen2.5-Max has recorded scores on 7 benchmarks — Lech Mazur Writing, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, FrontierMath-2025-02-28-Private, Aider polyglot, and 1 more. The full table above shows each score with its evidence status.
4 of 7 recorded scores (57%) are independently reproduced rather than self-reported by the lab.
Qwen2.5-Max has 7 tracked claims: 4 independently reproduced, 3 unverified.