Qwen·released 2025-07-301 source
Qwen3-30B-A3B-Thinking (Jul 2025) benchmark scores: 5 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 5 benchmarks·best result 70.25 on OTIS Mock AIME 2024-2025 (unverified)·0 of 5 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 70.25accuracy (%) | unverifiedT2 | 2025-07-30 |
| GPQA diamond | 60.1accuracy (%) | unverifiedT2 | 2025-07-30 |
| DTBench | 48.88accuracy (%) | unverifiedT1 | 2025-07-30https://conceptualreasoning.ai/dtbench |
| LMCA | 22.98accuracy (%) | unverifiedT1 | 2025-07-30https://conceptualreasoning.ai/lmca |
| Chess Puzzles | 3.2accuracy (%) | unverifiedT2 | 2025-07-30 |
Claims drawn from cited facts, not live model generation.
This model originates from Qwen, the developer behind it. It dates to 2025, a recent vintage for a model of its kind.
2 cited facts
The model is scored on 5 tracked benchmarks, but it holds no current top score on any of them. Its benchmark record is unverified: outside evaluation has neither confirmed nor disputed the results.
3 cited facts
Across its benchmarks, this model trails the state of the art by an average of 46.26 points, a gap wide enough to read as well behind the leaders rather than at the frontier or merely competitive. Consistent with that deficit, it holds the top score on none of its benchmarks (sota leader count of 0), so there is no category leadership to point to here.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Qwen3-30B-A3B-Thinking (Jul 2025) is an AI model developed by Qwen, released 2025-07-30. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Qwen3-30B-A3B-Thinking (Jul 2025) has recorded scores on 5 benchmarks, each shown with its evidence status.
Qwen3-30B-A3B-Thinking (Jul 2025) has recorded scores on 5 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, DTBench, LMCA, Chess Puzzles. The full table above shows each score with its evidence status.
0 of 5 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Qwen3-30B-A3B-Thinking (Jul 2025) has 5 tracked claims: 5 unverified.