Meta AI·released 2024-04-181 source
Llama 3-8B benchmark scores: 9 benchmarks tracked. 33% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 9 benchmarks·best result 77.07 on ARC AI2 (self-reported)·3 of 9 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| ARC AI2 | 77.07accuracy (%) | self-reported· optimizedT1 | 2024-04-18Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| OpenBookQA | 76.8accuracy (%) | self-reported· optimizedT1 | 2024-04-18Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| TriviaQA | 67.7accuracy (%) | self-reported· optimizedT1 | 2024-04-18Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| MMLU | 58.4accuracy (%) | unverified· optimizedT1 | 2024-04-18Stanford CRFM Leaderboard |
| Winogrande | 51.4accuracy (%) | unverified· optimizedT1 | 2024-04-18The Llama 3 Herd of Models |
| ANLI | 35.95accuracy (%) | self-reported· optimizedT1 | 2024-04-18Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| MATH level 5 | 6.13accuracy (%) | reproduced· optimizedT1 | 2024-04-18Epoch AI |
| GPQA diamond | 1.43accuracy (%) | reproduced· optimizedT1 | 2024-04-18Epoch AI |
| OTIS Mock AIME 2024-2025 | 0.73accuracy (%) | reproduced· optimizedT1 | 2024-04-18Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by Meta AI and released in 2024.
2 cited facts
At 9 tracked benchmarks, coverage is fair — enough to see a shape, not enough to end arguments. Reading the absent leads correctly matters: they rule out current leadership and rule on nothing else. Because its claim status is independently reproduced, outside evaluation backs up the scores, so the record can be trusted.
3 cited facts
With an average gap of 42.26 points behind the leader across benchmarks, the model falls into the wide gap band, indicating it is well behind the state-of-the-art.
1 cited fact
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Llama 3-8B is an AI model developed by Meta AI, released 2024-04-18. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Llama 3-8B has recorded scores on 9 benchmarks, each shown with its evidence status.
Llama 3-8B has recorded scores on 9 benchmarks — OpenBookQA, Winogrande, TriviaQA, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, and 3 more. The full table above shows each score with its evidence status.
3 of 9 recorded scores (33%) are independently reproduced rather than self-reported by the lab.
Llama 3-8B has 9 tracked claims: 3 independently reproduced, 4 self-reported, 2 unverified.