Qwen·released 2025-07-251 source
Qwen3-235B-A22B-Instruct (Jul 2025) benchmark scores: 5 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 5 benchmarks·best result 59.6 on Aider polyglot (unverified)·0 of 5 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| Aider polyglot | 59.6accuracy (%) | unverified· optimizedT1 | 2025-07-25Aider LLM Leaderboards |
| Fiction.LiveBench | 52.9accuracy (%) | unverifiedT1 | 2025-07-25Fiction.live leaderboard |
| WeirdML | 38.7accuracy (%) | unverifiedT1 | 2025-07-25WeirdML Leaderboard |
| ARC-AGI | 11accuracy (%) | unverified· optimizedT2 | 2025-07-25 |
| ARC-AGI-2 | 1.25accuracy (%) | unverifiedT2 | 2025-07-25 |
Claims drawn from cited facts, not live model generation.
This model was developed by Qwen and released in 2025.
2 cited facts
Only 5 tracked benchmarks inform this record, so extend any conclusion from it cautiously — the sample is thin. Holding no lead in that set is placement data; by itself it neither indicts nor bounds the model. Because its claim status is unverified, the reported numbers are neither confirmed nor disputed by outside evaluation.
3 cited facts
The model trails the state-of-the-art by an average of 58.52 points on its benchmarks, placing it well behind the leaders. It does not hold the top score on any benchmark, consistent with being far from the frontier.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Qwen3-235B-A22B-Instruct (Jul 2025) is an AI model developed by Qwen, released 2025-07-25. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Qwen3-235B-A22B-Instruct (Jul 2025) has recorded scores on 5 benchmarks, each shown with its evidence status.
Qwen3-235B-A22B-Instruct (Jul 2025) has recorded scores on 5 benchmarks — WeirdML, Aider polyglot, ARC-AGI, Fiction.LiveBench, ARC-AGI-2. The full table above shows each score with its evidence status.
0 of 5 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Qwen3-235B-A22B-Instruct (Jul 2025) has 5 tracked claims: 5 unverified.