Meta AI·released 2024-07-14✓ 2 sources
Llama 3.1-8B benchmark scores: 9 benchmarks tracked. 33% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 9 benchmarks·best result 82.4 on GSM8K (self-reported)·3 of 9 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| GSM8K | 82.4accuracy (%) | self-reported· optimizedT1 | 2024-07-23Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| PIQA | 62.4accuracy (%) | self-reported· optimizedT1 | 2024-07-23Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| MMLU | 41.47accuracy (%) | unverified· optimizedT1 | 2024-07-23Stanford CRFM Leaderboard |
| MATH level 5 | 22.88accuracy (%) | reproduced· optimizedT1 | 2024-07-23Epoch AI |
| Balrog | 15.1accuracy (%) | unverifiedT1 | 2024-07-23Balrog Leaderboard |
| OTIS Mock AIME 2024-2025 | 2.4accuracy (%) | reproduced· optimizedT1 | 2024-07-23Epoch AI |
| WeirdML | 1.73accuracy (%) | unverifiedT1 | 2024-07-23WeirdML Leaderboard |
| GPQA diamond | 1.26accuracy (%) | reproduced· optimizedT1 | 2024-07-23Epoch AI |
| CritPt | 0accuracy (%) | unverifiedT2 | 2024-07-23 |
Claims drawn from cited facts, not live model generation.
Developed by Meta AI, this model was released in 2024. With 8.03 billion parameters, it is a compact model, favoring efficiency and lower operating cost while offering less headroom than frontier-scale systems.
3 cited facts
This model is scored on nine tracked benchmarks. It holds no current top score among them. Because those results are independently reproduced, the record is externally backed and can be trusted as verified.
3 cited facts
Based on the disclosed harness, the model trails the leader by an average of 55.46 points, a wide gap that places it well behind the front-of-pack. It holds the top score on none of the tracked benchmarks, meaning it has no category leadership yet.
2 cited facts
2 facts cross-checked across data sources: 1 corroborated, 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, huggingface_models
Llama 3.1-8B is an AI model developed by Meta AI, released 2024-07-14. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Llama 3.1-8B has recorded scores on 9 benchmarks, each shown with its evidence status.
Llama 3.1-8B has recorded scores on 9 benchmarks — PIQA, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, MMLU, and 3 more. The full table above shows each score with its evidence status.
3 of 9 recorded scores (33%) are independently reproduced rather than self-reported by the lab.
Llama 3.1-8B has 9 tracked claims: 3 independently reproduced, 2 self-reported, 4 unverified.