Qwen·released 2025-04-281 source
Qwen3-8B benchmark scores: 6 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 6 benchmarks·best result 62.1 on Fiction.LiveBench (unverified)·0 of 6 independently reproduced·$0.18/$0.7 per M tokens
Source: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| Fiction.LiveBench | 62.1accuracy (%) | unverifiedT1 | 2025-04-28Fiction.live leaderboard |
| OTIS Mock AIME 2024-2025 | 56.07accuracy (%) | unverifiedT2 | 2025-04-28 |
| GPQA diamond | 42.34accuracy (%) | unverifiedT2 | 2025-04-28 |
| DTBench | 32.88accuracy (%) | unverifiedT1 | 2025-04-28https://conceptualreasoning.ai/dtbench |
| LMCA | 10.38accuracy (%) | unverifiedT1 | 2025-04-28https://conceptualreasoning.ai/lmca |
| Chess Puzzles | 0.04accuracy (%) | unverifiedT2 | 2025-04-28 |
Claims drawn from cited facts, not live model generation.
This model originates from Qwen, the developer that built it. It dates to 2025, making it a recent arrival in that developer's line of work.
2 cited facts
It is scored on 6 tracked benchmarks. It holds no current top score among them. Its record is unverified, meaning outside evaluation has neither confirmed nor disputed these numbers.
3 cited facts
On average, this model trails the leading score by 55.02 points, a wide gap that reads as well behind the leaders on its benchmarks. It holds the top score on 0 benchmarks, indicating no category leadership yet.
2 cited facts
5 facts cross-checked across data sources: 2 corroborated, 1 single-source, 2 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, modelsdev_models, openrouter_models
Qwen3-8B is an AI model developed by Qwen, released 2025-04-28. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Qwen3-8B has recorded scores on 6 benchmarks, each shown with its evidence status.
Qwen3-8B has recorded scores on 6 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, DTBench, LMCA, Chess Puzzles, Fiction.LiveBench. The full table above shows each score with its evidence status.
0 of 6 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Qwen3-8B has 6 tracked claims: 6 unverified.
Listed API pricing: $0.18 per million input tokens, $0.7 per million output tokens (prices disputed across sources). See the pricing block for the full breakdown.