Meta AI·released 2023-02-271 source
LLaMA-13B benchmark scores: 11 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 11 benchmarks·best result 77.9 on TriviaQA (unverified)·0 of 11 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| TriviaQA | 77.9accuracy (%) | unverified· optimizedT1 | 2023-02-27Llama 2: Open Foundation and Fine-Tuned Chat Models |
| LAMBADA | 75.2accuracy (%) | self-reported· optimizedT1 | 2023-02-27Qwen Technical Report |
| HellaSwag | 72.27accuracy (%) | unverified· optimizedT1 | 2023-02-27LLaMA: Open and Efficient Foundation Language Models |
| PIQA | 60.2accuracy (%) | self-reported· optimizedT1 | 2023-02-27Qwen Technical Report |
| Winogrande | 46accuracy (%) | unverified· optimizedT1 | 2023-02-27LLaMA: Open and Efficient Foundation Language Models |
| OpenBookQA | 41.87accuracy (%) | unverified· optimizedT1 | 2023-02-27Llama 2: Open Foundation and Fine-Tuned Chat Models |
| ARC AI2 | 36.93accuracy (%) | self-reported· optimizedT1 | 2023-02-27Qwen Technical Report |
| MMLU | 30.27accuracy (%) | unverified· optimizedT1 | 2023-02-27Llama 2: Open Foundation and Fine-Tuned Chat Models |
| ScienceQA | 24.44accuracy (%) | unverified· optimizedT1 | 2023-02-27ScienceQA leaderboard |
| GSM8K | 20.55accuracy (%) | unverified· optimizedT1 | 2023-02-27Stanford HELM |
| BBH | 17.2accuracy (%) | unverified· optimizedT1 | 2023-02-27Baichuan 2: Open Large-scale Language Models |
Claims drawn from cited facts, not live model generation.
This model was developed by Meta AI and released in 2023.
2 cited facts
Use the breadth: at 11 tracked scores, cross-benchmark patterns are steadier ground than any single result. Nothing at the front, which is a bound rather than a placement — the record shows where the model is not, rather than where it is. These scores are self-reported — vendor-claimed and not yet independently confirmed — so readers should treat them as claims rather than verified results.
3 cited facts
The model trails the state‑of‑the‑art by an average of 39.99 points, a wide gap that places it well behind the current leaders. It holds the top score on none of the tracked benchmarks, confirming it has no genuine category leadership.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
LLaMA-13B is an AI model developed by Meta AI, released 2023-02-27. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
LLaMA-13B has recorded scores on 11 benchmarks, each shown with its evidence status.
LLaMA-13B has recorded scores on 11 benchmarks — PIQA, OpenBookQA, Winogrande, TriviaQA, ScienceQA, MMLU, and 5 more. The full table above shows each score with its evidence status.
0 of 11 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
LLaMA-13B has 11 tracked claims: 3 self-reported, 8 unverified.