Meta AI·released 2023-07-181 source
Llama 2-34B benchmark scores: 8 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 8 benchmarks·best result 84.6 on TriviaQA (unverified)·0 of 8 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| TriviaQA | 84.6accuracy (%) | unverified· optimizedT1 | 2023-07-18Llama 2: Open Foundation and Fine-Tuned Chat Models |
| PIQA | 63.8accuracy (%) | unverified· optimizedT1 | 2023-07-18Llama 2: Open Foundation and Fine-Tuned Chat Models |
| Winogrande | 53.4accuracy (%) | unverified· optimizedT1 | 2023-07-18Llama 2: Open Foundation and Fine-Tuned Chat Models |
| MMLU | 50.13accuracy (%) | unverified· optimizedT1 | 2023-07-18Llama 2: Open Foundation and Fine-Tuned Chat Models |
| OpenBookQA | 44.27accuracy (%) | unverified· optimizedT1 | 2023-07-18Llama 2: Open Foundation and Fine-Tuned Chat Models |
| GSM8K | 42.2accuracy (%) | self-reported· optimizedT1 | 2023-07-18Nemotron-4 15B Technical Report |
| ARC AI2 | 39.33accuracy (%) | unverified· optimizedT1 | 2023-07-18The Falcon Series of Open Language Models |
| BBH | 25.47accuracy (%) | self-reported· optimizedT1 | 2023-07-18Nemotron-4 15B Technical Report |
Claims drawn from cited facts, not live model generation.
This model was developed by Meta AI and released in 2023.
2 cited facts
Context rather than verdict: 8 scored benchmarks establish coverage, while the absent lead bounds the record's ceiling without grading the model. The record is self-reported, meaning the numbers are vendor-claimed and not independently confirmed; therefore, readers should treat them as claims rather than verified results.
3 cited facts
This model trails the best score by an average of 35.17 points, a wide gap that places it well behind the leaders. It does not hold the top score on any benchmark, meaning it is not a genuine category leader.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Llama 2-34B is an AI model developed by Meta AI, released 2023-07-18. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Llama 2-34B has recorded scores on 8 benchmarks, each shown with its evidence status.
Llama 2-34B has recorded scores on 8 benchmarks — PIQA, OpenBookQA, Winogrande, TriviaQA, MMLU, ARC AI2, and 2 more. The full table above shows each score with its evidence status.
0 of 8 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Llama 2-34B has 8 tracked claims: 2 self-reported, 6 unverified.