Qwen·released 2024-09-191 source
Qwen2.5-7B benchmark scores: 7 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 7 benchmarks·best result 63.87 on MMLU (unverified)·0 of 7 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MMLU | 63.87accuracy (%) | unverifiedT1 | 2024-09-19Stanford CRFM Leaderboard |
| GPQA diamond | 13.97accuracy (%) | unverifiedT2 | 2024-09-19 |
| DTBench | 12.88accuracy (%) | unverifiedT1 | 2024-09-19https://conceptualreasoning.ai/dtbench |
| Balrog | 7.8accuracy (%) | unverifiedT1 | 2024-09-19https://balrogai.com/ |
| LMCA | 7.55accuracy (%) | unverifiedT1 | 2024-09-19https://conceptualreasoning.ai/lmca |
| OTIS Mock AIME 2024-2025 | 2.4accuracy (%) | unverifiedT2 | 2024-09-19 |
| Chess Puzzles | 0accuracy (%) | unverifiedT2 | 2024-09-19 |
Claims drawn from cited facts, not live model generation.
This model was developed by Qwen and released in 2024.
2 cited facts
This model is scored on 7 tracked benchmarks, but it currently holds no top score on any of them. Its record is unverified: no outside evaluation has confirmed or disputed the reported results.
3 cited facts
On average this model trails the benchmark leader by 67.21 points under the disclosed harness, and a gap of that width reads as well behind the front of the pack rather than at or near the frontier. Consistently, it holds the top score on 0 benchmarks, so there is no category leadership here to offset that deficit.
3 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Qwen2.5-7B is an AI model developed by Qwen, released 2024-09-19. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Qwen2.5-7B has recorded scores on 7 benchmarks, each shown with its evidence status.
Qwen2.5-7B has recorded scores on 7 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, DTBench, MMLU, LMCA, Chess Puzzles, and 1 more. The full table above shows each score with its evidence status.
0 of 7 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Qwen2.5-7B has 7 tracked claims: 7 unverified.