Qwen·released 2025-04-27✓ 2 sources
Qwen3-235B-A22B benchmark scores: 7 benchmarks tracked. 29% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 7 benchmarks·best result 83 on Lech Mazur Writing (unverified)·2 of 7 independently reproduced·$0.2/$0.6 per M tokens
Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| Lech Mazur Writing | 83accuracy (%) | unverifiedT1 | 2025-04-28lechmazur/writing Github repository |
| MATH level 5 | 68.86accuracy (%) | reproduced· optimizedT1 | 2025-04-28Epoch AI |
| Fiction.LiveBench | 67.7accuracy (%) | unverifiedT1 | 2025-04-28Fiction.live leaderboard |
| GPQA diamond | 60.94accuracy (%) | reproduced· optimizedT1 | 2025-04-28Epoch AI |
| Aider polyglot | 59.6accuracy (%) | unverified· optimizedT1 | 2025-04-28Aider LLM Leaderboards |
| WeirdML | 37.28accuracy (%) | unverifiedT1 | 2025-04-28WeirdML Leaderboard |
| SimpleBench | 17.2accuracy (%) | unverifiedT1 | 2025-04-28SimpleBench Leaderboard |
Claims drawn from cited facts, not live model generation.
Developed by Qwen, this model was released in 2025 and has approximately 235 billion parameters, placing it at frontier-scale, which implies greater headroom but also higher computational cost.
3 cited facts
This model is scored on 7 tracked benchmarks and holds no current top score; because its results have been independently reproduced, the scores are verified by outside evaluation.
3 cited facts
The model trails the state-of-the-art by an average of 33.51 points, a wide gap that places it well behind the leaders. It does not hold the top score on any evaluated benchmark, indicating no genuine category leadership.
2 cited facts
6 facts cross-checked across data sources: 1 corroborated, 1 single-source, 4 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, huggingface_models, litellm_prices, modelsdev_models +1 more
Qwen3-235B-A22B is an AI model developed by Qwen, released 2025-04-27. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Qwen3-235B-A22B has recorded scores on 7 benchmarks, each shown with its evidence status.
Qwen3-235B-A22B has recorded scores on 7 benchmarks — Lech Mazur Writing, GPQA diamond, MATH level 5, WeirdML, SimpleBench, Aider polyglot, and 1 more. The full table above shows each score with its evidence status.
2 of 7 recorded scores (29%) are independently reproduced rather than self-reported by the lab.
Qwen3-235B-A22B has 7 tracked claims: 2 independently reproduced, 5 unverified.
Listed API pricing: $0.2 per million input tokens, $0.6 per million output tokens (prices disputed across sources). See the pricing block for the full breakdown.