Meta AI·released 2024-07-231 source
Llama 3.1-70B benchmark scores: 7 benchmarks tracked. 43% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 7 benchmarks·best result 73.47 on MMLU (unverified)·3 of 7 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MMLU | 73.47accuracy (%) | unverified· optimizedT1 | 2024-07-23Stanford CRFM Leaderboard |
| MATH level 5 | 36.68accuracy (%) | reproduced· optimizedT1 | 2024-07-23Epoch AI |
| Balrog | 27.9accuracy (%) | unverifiedT1 | 2024-07-23Balrog Leaderboard |
| GPQA diamond | 25.59accuracy (%) | reproduced· optimizedT1 | 2024-07-23Epoch AI |
| WeirdML | 8.97accuracy (%) | unverifiedT1 | 2024-07-23WeirdML Leaderboard |
| The Agent Company | 6.9accuracy (%) | unverifiedT1 | 2024-07-23TheAgentCompany experiment results github |
| OTIS Mock AIME 2024-2025 | 3.51accuracy (%) | reproduced· optimizedT1 | 2024-07-23Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by Meta AI and released in 2024.
2 cited facts
Hold this record lightly: 7 scored benchmarks is thin coverage, and the missing lead adds a ceiling, not a grade. Because the record is independently reproduced, outside evaluation backs these scores up.
3 cited facts
The model trails the state-of-the-art by an average of 54.41 points, a wide gap that places it well behind the leaders, and it does not hold the top score on any benchmark, confirming it is not a category leader.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Llama 3.1-70B is an AI model developed by Meta AI, released 2024-07-23. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Llama 3.1-70B has recorded scores on 7 benchmarks, each shown with its evidence status.
Llama 3.1-70B has recorded scores on 7 benchmarks — GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, MMLU, Balrog, and 1 more. The full table above shows each score with its evidence status.
3 of 7 recorded scores (43%) are independently reproduced rather than self-reported by the lab.
Llama 3.1-70B has 7 tracked claims: 3 independently reproduced, 4 unverified.