Meta AI·released 2023-07-181 source
Llama 2-13B benchmark scores: 11 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 11 benchmarks·best result 79.6 on TriviaQA (unverified)·0 of 11 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| TriviaQA | 79.6accuracy (%) | unverified· optimizedT1 | 2023-07-18Gemma: Open Models Based on Gemini Research and Technology |
| LAMBADA | 76.5accuracy (%) | self-reported· optimizedT1 | 2023-07-18Qwen Technical Report |
| HellaSwag | 74.27accuracy (%) | unverified· optimizedT1 | 2023-07-18Gemma: Open Models Based on Gemini Research and Technology |
| PIQA | 61.6accuracy (%) | unverified· optimizedT1 | 2023-07-18Gemma: Open Models Based on Gemini Research and Technology |
| ARC AI2 | 47.07accuracy (%) | unverified· optimizedT1 | 2023-07-18The Falcon Series of Open Language Models |
| Winogrande | 45.6accuracy (%) | unverified· optimizedT1 | 2023-07-18Llama 2: Open Foundation and Fine-Tuned Chat Models |
| BBH | 44.27accuracy (%) | unverified· optimizedT1 | 2023-07-18Gemma: Open Models Based on Gemini Research and Technology |
| OpenBookQA | 42.67accuracy (%) | unverified· optimizedT1 | 2023-07-18Llama 2: Open Foundation and Fine-Tuned Chat Models |
| ScienceQA | 41.04accuracy (%) | unverified· optimizedT1 | 2023-07-18ScienceQA leaderboard |
| MMLU | 40.8accuracy (%) | unverified· optimizedT1 | 2023-07-18Stanford CRFM Leaderboard |
| GSM8K | 36.9accuracy (%) | unverified· optimizedT1 | 2023-07-18Stanford HELM |
Claims drawn from cited facts, not live model generation.
This model was developed by Meta AI and released in 2023.
2 cited facts
Its coverage spans 11 tracked benchmarks — wide enough that a stray result should not steer the overall read. Front of the table it is not, on any of them — a limit on leadership claims that implies nothing about the rest of its placements. Because its scores are self-reported without independent confirmation, they should be read as claims rather than verified results.
3 cited facts
This model trails the state-of-the-art by an average of 32.04 points across benchmarks, placing it well behind the leaders. It tops nothing in this set, and that is the full extent of what the figure establishes — context for reading the scores, not a conclusion about the model.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Llama 2-13B is an AI model developed by Meta AI, released 2023-07-18. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Llama 2-13B has recorded scores on 11 benchmarks, each shown with its evidence status.
Llama 2-13B has recorded scores on 11 benchmarks — PIQA, OpenBookQA, Winogrande, TriviaQA, ScienceQA, MMLU, and 5 more. The full table above shows each score with its evidence status.
0 of 11 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Llama 2-13B has 11 tracked claims: 1 self-reported, 10 unverified.