Meta AI·released 2023-07-13✓ 2 sources
Llama 2-7B benchmark scores: 11 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 11 benchmarks·best result 73.7 on TriviaQA (unverified)·0 of 11 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| TriviaQA | 73.7accuracy (%) | unverified· optimizedT1 | 2023-07-18Gemma: Open Models Based on Gemini Research and Technology |
| LAMBADA | 73.3accuracy (%) | self-reported· optimizedT1 | 2023-07-18Qwen Technical Report |
| HellaSwag | 69.6accuracy (%) | unverified· optimizedT1 | 2023-07-18Gemma: Open Models Based on Gemini Research and Technology |
| PIQA | 57.6accuracy (%) | unverified· optimizedT1 | 2023-07-18Gemma: Open Models Based on Gemini Research and Technology |
| OpenBookQA | 44.8accuracy (%) | unverified· optimizedT1 | 2023-07-18Llama 2: Open Foundation and Fine-Tuned Chat Models |
| Winogrande | 38.4accuracy (%) | unverified· optimizedT1 | 2023-07-18Llama 2: Open Foundation and Fine-Tuned Chat Models |
| ARC AI2 | 27.87accuracy (%) | unverified· optimizedT1 | 2023-07-18The Falcon Series of Open Language Models |
| MMLU | 27.73accuracy (%) | unverified· optimizedT1 | 2023-07-18Stanford CRFM Leaderboard |
| ScienceQA | 24.11accuracy (%) | unverified· optimizedT1 | 2023-07-18ScienceQA leaderboard |
| BBH | 18.88accuracy (%) | unverified· optimizedT1 | 2023-07-18Gemma: Open Models Based on Gemini Research and Technology |
| GSM8K | 16.7accuracy (%) | unverified· optimizedT1 | 2023-07-18Gemma: Open Models Based on Gemini Research and Technology |
Claims drawn from cited facts, not live model generation.
This model, developed by Meta AI and released in 2023, is a compact 6.74-billion-parameter model, making it efficient and cost-effective compared to larger alternatives.
3 cited facts
The model is scored on 11 tracked benchmarks, but holds no current top score; these numbers are self-reported, so they should be treated as claims rather than verified results.
3 cited facts
With an average gap of 42.73 points to the state-of-the-art score on its benchmarks, this model falls well behind the leaders, indicating a substantial performance deficit. It does not hold the top score on any benchmark, consistent with the wide gap.
2 cited facts
2 facts cross-checked across data sources: 1 corroborated, 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, huggingface_models
Llama 2-7B is an AI model developed by Meta AI, released 2023-07-13. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Llama 2-7B has recorded scores on 11 benchmarks, each shown with its evidence status.
Llama 2-7B has recorded scores on 11 benchmarks — PIQA, OpenBookQA, Winogrande, TriviaQA, ScienceQA, MMLU, and 5 more. The full table above shows each score with its evidence status.
0 of 11 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Llama 2-7B has 11 tracked claims: 1 self-reported, 10 unverified.