Meta AI·released 2023-02-241 source
LLaMA-7B benchmark scores: 11 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 11 benchmarks·best result 73.3 on LAMBADA (self-reported)·0 of 11 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| LAMBADA | 73.3accuracy (%) | self-reported· optimizedT1 | 2023-02-24Qwen Technical Report |
| TriviaQA | 71accuracy (%) | unverified· optimizedT1 | 2023-02-24Llama 2: Open Foundation and Fine-Tuned Chat Models |
| HellaSwag | 68.27accuracy (%) | unverified· optimizedT1 | 2023-02-24LLaMA: Open and Efficient Foundation Language Models |
| PIQA | 59.6accuracy (%) | self-reported· optimizedT1 | 2023-02-24Qwen Technical Report |
| OpenBookQA | 42.93accuracy (%) | unverified· optimizedT1 | 2023-02-24Llama 2: Open Foundation and Fine-Tuned Chat Models |
| Winogrande | 40.2accuracy (%) | unverified· optimizedT1 | 2023-02-24LLaMA: Open and Efficient Foundation Language Models |
| ARC AI2 | 30.13accuracy (%) | self-reported· optimizedT1 | 2023-02-24Qwen Technical Report |
| ScienceQA | 14.92accuracy (%) | unverified· optimizedT1 | 2023-02-24ScienceQA leaderboard |
| MMLU | 14.13accuracy (%) | unverified· optimizedT1 | 2023-02-24Llama 2: Open Foundation and Fine-Tuned Chat Models |
| BBH | 11.33accuracy (%) | unverified· optimizedT1 | 2023-02-24Baichuan 2: Open Large-scale Language Models |
| GSM8K | 11accuracy (%) | unverified· optimizedT1 | 2023-02-24Stanford HELM |
Claims drawn from cited facts, not live model generation.
This model was developed by Meta AI and released in 2023.
2 cited facts
This model is tracked on 11 benchmarks but holds no top score; because the results are self-reported, they are vendor-claimed and have not been independently reproduced, so they should be interpreted as claims rather than verified findings.
3 cited facts
With an average SOTA gap of 46.0 points, the model lags well behind the front-of-pack leaders.
1 cited fact
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
LLaMA-7B is an AI model developed by Meta AI, released 2023-02-24. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
LLaMA-7B has recorded scores on 11 benchmarks, each shown with its evidence status.
LLaMA-7B has recorded scores on 11 benchmarks — PIQA, OpenBookQA, Winogrande, TriviaQA, ScienceQA, MMLU, and 5 more. The full table above shows each score with its evidence status.
0 of 11 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
LLaMA-7B has 11 tracked claims: 3 self-reported, 8 unverified.