Meta AI·released 2024-07-231 source
Llama 3.1-405B benchmark scores: 14 benchmarks tracked, leading 2. 21% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 2 of 14 benchmarks·best result 93.73 on ARC AI2 (self-reported)·3 of 14 independently reproduced
Head-to-headLlama 3.1-405B vs Qwen2.5-72BLeads on: ARC AI2, Winogrande
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| ARC AI2 | 93.73accuracy (%) | self-reported· optimizedT1 | 2024-07-23DeepSeek-V3 Technical Report |
| HellaSwag | 85.6accuracy (%) | self-reported· optimizedT1 | 2024-07-23DeepSeek-V3 Technical Report |
| TriviaQA | 82.7accuracy (%) | self-reported· optimizedT1 | 2024-07-23DeepSeek-V3 Technical Report |
| MMLU | 79.33accuracy (%) | unverified· optimizedT1 | 2024-07-23Stanford CRFM Leaderboard |
| Winogrande | 78.4accuracy (%) | unverified· optimizedT1 | 2024-07-23The Llama 3 Herd of Models |
| BBH | 77.2accuracy (%) | self-reported· optimizedT1 | 2024-07-23DeepSeek-V3 Technical Report |
| PIQA | 71.8accuracy (%) | self-reported· optimizedT1 | 2024-07-23DeepSeek-V3 Technical Report |
| MATH level 5 | 49.77accuracy (%) | reproduced· optimizedT1 | 2024-07-23Epoch AI |
| GPQA diamond | 34.55accuracy (%) | reproduced· optimizedT1 | 2024-07-23Epoch AI |
| WeirdML | 21.38accuracy (%) | unverifiedT1 | 2024-07-23WeirdML Leaderboard |
| OTIS Mock AIME 2024-2025 | 9.63accuracy (%) | reproduced· optimizedT1 | 2024-07-23Epoch AI |
| SimpleBench | 7.6accuracy (%) | unverifiedT1 | 2024-07-23SimpleBench Leaderboard |
| Cybench | 7.5accuracy (%) | unverified· optimizedT1 | 2024-07-23Cybench leaderboard |
| The Agent Company | 7.4accuracy (%) | unverifiedT1 | 2024-07-23TheAgentCompany experiment results github |
Claims drawn from cited facts, not live model generation.
This model was developed by Meta AI and released in 2024.
2 cited facts
This model is scored on 14 tracked benchmarks and holds current top scores on 2 of them. Because its scores have been independently reproduced, outside evaluation backs these numbers up.
3 cited facts
The model trails the state-of-the-art by an average gap of 34.9 points, placing it well behind the leaders. It holds the top score on two benchmarks, indicating pockets of leadership.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: curated_capability_claims, epoch_benchmarks
Llama 3.1-405B is an AI model developed by Meta AI, released 2024-07-23. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Llama 3.1-405B has recorded scores on 14 benchmarks, each shown with its evidence status.
Llama 3.1-405B has recorded scores on 14 benchmarks — PIQA, Winogrande, TriviaQA, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, and 8 more. The full table above shows each score with its evidence status.
3 of 14 recorded scores (21%) are independently reproduced rather than self-reported by the lab.
Llama 3.1-405B has 14 tracked claims: 3 independently reproduced, 5 self-reported, 6 unverified.