Qwen·released 2024-06-071 source
Qwen2-72B benchmark scores: 6 benchmarks tracked. 33% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 6 benchmarks·best result 76.53 on MMLU (unverified)·2 of 6 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MMLU | 76.53accuracy (%) | unverified· optimizedT1 | 2024-06-07Stanford CRFM Leaderboard |
| MATH level 5 | 39.07accuracy (%) | reproduced· optimizedT1 | 2024-06-07Epoch AI |
| METR Time Horizons | 29.9accuracy (%) | unverifiedT1 | 2024-06-07METR - Measuring AI Ability to Complete Long Tasks |
| GPQA diamond | 21.04accuracy (%) | reproduced· optimizedT1 | 2024-06-07Epoch AI |
| WeirdML | 11.3accuracy (%) | unverifiedT1 | 2024-06-07https://htihle.github.io/weirdml.html |
| The Agent Company | 1.1accuracy (%) | unverifiedT1 | 2024-06-07TheAgentCompany experiment results github |
Claims drawn from cited facts, not live model generation.
This model was developed by Qwen and released in 2024.
2 cited facts
This model is tracked across six benchmarks. It holds no current top score among them. Because the record has been independently reproduced, these results are verified rather than vendor-claimed.
3 cited facts
With an average gap of 51.69 points behind the best score, this model sits well behind the leaders on its benchmarks — a wide gap, not a competitive or frontier position. It holds the top score on none of the tracked benchmarks, meaning it has no category leadership yet.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Qwen2-72B is an AI model developed by Qwen, released 2024-06-07. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Qwen2-72B has recorded scores on 6 benchmarks, each shown with its evidence status.
Qwen2-72B has recorded scores on 6 benchmarks — GPQA diamond, MATH level 5, WeirdML, MMLU, METR Time Horizons, The Agent Company. The full table above shows each score with its evidence status.
2 of 6 recorded scores (33%) are independently reproduced rather than self-reported by the lab.
Qwen2-72B has 6 tracked claims: 2 independently reproduced, 4 unverified.