Qwen·released 2024-09-191 source
Qwen2.5-72B benchmark scores: 16 benchmarks tracked. 19% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 16 benchmarks·best result 92.67 on ARC AI2 (self-reported)·3 of 16 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| ARC AI2 | 92.67accuracy (%) | self-reported· optimizedT1 | 2024-09-19DeepSeek-V3 Technical Report |
| MMLU | 80.4accuracy (%) | self-reported· optimizedT1 | 2024-09-19DeepSeek-V3 Technical Report |
| HellaSwag | 79.73accuracy (%) | self-reported· optimizedT1 | 2024-09-19DeepSeek-V3 Technical Report |
| BBH | 73.07accuracy (%) | self-reported· optimizedT1 | 2024-09-19DeepSeek-V3 Technical Report |
| TriviaQA | 71.9accuracy (%) | self-reported· optimizedT1 | 2024-09-19DeepSeek-V3 Technical Report |
| PIQA | 65.2accuracy (%) | self-reported· optimizedT1 | 2024-09-19DeepSeek-V3 Technical Report |
| Winogrande | 64.6accuracy (%) | self-reported· optimizedT1 | 2024-09-19DeepSeek-V3 Technical Report |
| MATH level 5 | 63.17accuracy (%) | reproduced· optimizedT1 | 2024-09-19Epoch AI |
| GeoBench | 62accuracy (%) | unverifiedT1 | 2024-09-19GeoBench leaderboard |
| METR Time Horizons | 35.78accuracy (%) | unverifiedT1 | 2024-09-19METR - Measuring AI Ability to Complete Long Tasks |
| GPQA diamond | 32.2accuracy (%) | reproduced· optimizedT1 | 2024-09-19Epoch AI |
| Balrog | 16.2accuracy (%) | unverifiedT1 | 2024-09-19Balrog Leaderboard |
| WeirdML | 15.97accuracy (%) | unverifiedT1 | 2024-09-19WeirdML Leaderboard |
| OTIS Mock AIME 2024-2025 | 7.96accuracy (%) | reproduced· optimizedT1 | 2024-09-19Epoch AI |
| The Agent Company | 5.7accuracy (%) | unverifiedT1 | 2024-09-19TheAgentCompany experiment results github |
| OSWorld | 5accuracy (%) | unverified· optimizedT1 | 2024-09-19OS World Website |
Claims drawn from cited facts, not live model generation.
This model, developed by Qwen, was released in 2024.
2 cited facts
This model is scored on 16 tracked benchmarks, but it currently holds no top score on any of them. Because the record has been independently reproduced, these results are backed by outside evaluation and can be treated as verified.
3 cited facts
The model trails the benchmark leader by an average of 34.62 points, which is a wide gap and clearly well behind the front of the pack. It holds the top score on none of the tracked benchmarks, so there is no category leadership to claim.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Qwen2.5-72B is an AI model developed by Qwen, released 2024-09-19. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Qwen2.5-72B has recorded scores on 16 benchmarks, each shown with its evidence status.
Qwen2.5-72B has recorded scores on 16 benchmarks — PIQA, Winogrande, TriviaQA, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, and 10 more. The full table above shows each score with its evidence status.
3 of 16 recorded scores (19%) are independently reproduced rather than self-reported by the lab.
Qwen2.5-72B has 16 tracked claims: 3 independently reproduced, 7 self-reported, 6 unverified.