Qwen·released 2025-07-251 source
Qwen3-235B-A22B-Thinking (Jul 2025) benchmark scores: 10 benchmarks tracked. 60% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 10 benchmarks·best result 86.65 on OTIS Mock AIME 2024-2025 (reproduced)·6 of 10 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 86.65accuracy (%) | reproduced· optimizedT1 | 2025-07-25Epoch AI |
| Lech Mazur Writing | 82.4accuracy (%) | unverifiedT1 | 2025-07-25lechmazur/writing Github repository |
| Fiction.LiveBench | 75accuracy (%) | unverifiedT1 | 2025-07-25Fiction.live leaderboard |
| GPQA diamond | 73.4accuracy (%) | reproduced· optimizedT1 | 2025-07-25Epoch AI |
| SimpleQA Verified | 50.1accuracy (%) | reproducedT1 | 2025-07-25Epoch AI |
| WeirdML | 41.04accuracy (%) | unverifiedT1 | 2025-07-25WeirdML Leaderboard |
| FrontierMath-2025-02-28-Private | 14.88accuracy (%) | reproduced· optimizedT1 | 2025-07-25Epoch AI |
| Chess Puzzles | 7.41accuracy (%) | reproducedT1 | 2025-07-25Epoch AI |
| CritPt | 0accuracy (%) | unverifiedT2 | 2025-07-25 |
| FrontierMath-Tier-4-2025-07-01-Private | 0accuracy (%) | reproducedT1 | 2025-07-25Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by Qwen and released in 2025. As a 2025 release it is a recent model, though its scale remains unspecified here.
3 cited facts
This model is scored on 10 tracked benchmarks. It currently holds no top score on any of them. The scores are backed by independent reproduction, so the record is reliable.
3 cited facts
On average, this model trails the leader by 30.88 points on its benchmarks, a wide gap that places it well behind the front of the pack. It holds the top score on none of the benchmarks, so it currently has no category leadership.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Qwen3-235B-A22B-Thinking (Jul 2025) is an AI model developed by Qwen, released 2025-07-25. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Qwen3-235B-A22B-Thinking (Jul 2025) has recorded scores on 10 benchmarks, each shown with its evidence status.
Qwen3-235B-A22B-Thinking (Jul 2025) has recorded scores on 10 benchmarks — Lech Mazur Writing, GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, FrontierMath-2025-02-28-Private, and 4 more. The full table above shows each score with its evidence status.
6 of 10 recorded scores (60%) are independently reproduced rather than self-reported by the lab.
Qwen3-235B-A22B-Thinking (Jul 2025) has 10 tracked claims: 6 independently reproduced, 4 unverified.