Qwen·released 2025-04-291 source
Qwen3-32B benchmark scores: 7 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 7 benchmarks·best result 74.2 on Fiction.LiveBench (unverified)·0 of 7 independently reproduced·$0.7/$2.8 per M tokens
Source: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| Fiction.LiveBench | 74.2accuracy (%) | unverifiedT1 | 2025-04-28Fiction.live leaderboard |
| OTIS Mock AIME 2024-2025 | 66.91accuracy (%) | unverifiedT2 | 2025-04-29 |
| GPQA diamond | 54.29accuracy (%) | unverifiedT2 | 2025-04-29 |
| DTBench | 45.78accuracy (%) | unverifiedT1 | 2025-04-28https://conceptualreasoning.ai/dtbench |
| Aider polyglot | 40accuracy (%) | unverifiedT1 | 2025-04-28https://aider.chat/docs/leaderboards/ |
| LMCA | 20.33accuracy (%) | unverifiedT1 | 2025-04-28https://conceptualreasoning.ai/lmca |
| Chess Puzzles | 0.04accuracy (%) | unverifiedT2 | 2025-04-29 |
Claims drawn from cited facts, not live model generation.
Developed by Qwen, this model was released in 2025.
2 cited facts
This model is scored on 7 tracked benchmarks and currently tops none of them. Its record is unverified: outside evaluation has neither confirmed nor disputed the results, so they should be read as claims rather than verified outcomes.
3 cited facts
On the disclosed harness, this model trails the benchmark leader by an average of 45.76 points, a wide gap that places it well behind the front of the pack rather than merely a couple of points off the frontier. Consistent with that, it holds the top score on 0 of its benchmarks, meaning there is no category where it currently leads — no genuine category leadership yet, though the gap is a measured score shortfall under a stated harness rather than a claim about capability.
3 cited facts
5 facts cross-checked across data sources: 2 corroborated, 1 single-source, 2 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, modelsdev_models, openrouter_models
Qwen3-32B is an AI model developed by Qwen, released 2025-04-29. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Qwen3-32B has recorded scores on 7 benchmarks, each shown with its evidence status.
Qwen3-32B has recorded scores on 7 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, DTBench, LMCA, Chess Puzzles, Aider polyglot, and 1 more. The full table above shows each score with its evidence status.
0 of 7 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Qwen3-32B has 7 tracked claims: 7 unverified.
Listed API pricing: $0.7 per million input tokens, $2.8 per million output tokens (prices disputed across sources). See the pricing block for the full breakdown.