Meta AI·released 2023-07-181 source
Llama 2-70B benchmark scores: 13 benchmarks tracked, leading 1. 23% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 1 of 13 benchmarks·best result 87.6 on TriviaQA (unverified)·3 of 13 independently reproduced
Head-to-headLlama 2-70B vs Claude 2Leads on: TriviaQA
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| TriviaQA | 87.6accuracy (%) | unverified· optimizedT1 | 2023-07-18Llama 2: Open Foundation and Fine-Tuned Chat Models |
| HellaSwag | 80.4accuracy (%) | self-reported· optimizedT1 | 2023-07-18Qwen Technical Report |
| LAMBADA | 78.9accuracy (%) | self-reported· optimizedT1 | 2023-07-18Qwen Technical Report |
| ARC AI2 | 71.07accuracy (%) | unverified· optimizedT1 | 2023-07-18The Falcon Series of Open Language Models |
| GSM8K | 69.6accuracy (%) | unverified· optimizedT1 | 2023-07-18Stanford HELM |
| PIQA | 65.6accuracy (%) | self-reported· optimizedT1 | 2023-07-18Qwen Technical Report |
| Winogrande | 60.4accuracy (%) | unverified· optimizedT1 | 2023-07-18Llama 2: Open Foundation and Fine-Tuned Chat Models |
| MMLU | 59.87accuracy (%) | unverified· optimizedT1 | 2023-07-18Stanford CRFM Leaderboard |
| BBH | 53.2accuracy (%) | self-reported· optimizedT1 | 2023-07-18Qwen Technical Report |
| OpenBookQA | 46.93accuracy (%) | unverified· optimizedT1 | 2023-07-18Llama 2: Open Foundation and Fine-Tuned Chat Models |
| MATH level 5 | 3.29accuracy (%) | reproduced· optimizedT1 | 2023-07-18Epoch AI |
| GPQA diamond | 1.77accuracy (%) | reproduced· optimizedT1 | 2023-07-18Epoch AI |
| OTIS Mock AIME 2024-2025 | 0accuracy (%) | reproduced· optimizedT1 | 2023-07-18Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by Meta AI and released in 2023.
2 cited facts
That 1 lead among 13 scored benchmarks is the cell to open first — a lone top score is a fact about a single table's conditions until the rest of the record corroborates it. These scores are independently reproduced, meaning outside evaluation backs them up.
3 cited facts
This model trails the state-of-the-art by an average gap of 36.18 points, placing it well behind the leading models, although it holds the top score on 1 benchmark.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Llama 2-70B is an AI model developed by Meta AI, released 2023-07-18. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Llama 2-70B has recorded scores on 13 benchmarks, each shown with its evidence status.
Llama 2-70B has recorded scores on 13 benchmarks — PIQA, OpenBookQA, Winogrande, TriviaQA, GPQA diamond, MATH level 5, and 7 more. The full table above shows each score with its evidence status.
3 of 13 recorded scores (23%) are independently reproduced rather than self-reported by the lab.
Llama 2-70B has 13 tracked claims: 3 independently reproduced, 4 self-reported, 6 unverified.