Meta AI·released 2023-02-241 source
LLaMA-65B benchmark scores: 10 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 10 benchmarks·best result 86 on TriviaQA (unverified)·0 of 10 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| TriviaQA | 86accuracy (%) | unverified· optimizedT1 | 2023-02-24Llama 2: Open Foundation and Fine-Tuned Chat Models |
| HellaSwag | 78.93accuracy (%) | unverified· optimizedT1 | 2023-02-24LLaMA: Open and Efficient Foundation Language Models |
| LAMBADA | 77.7accuracy (%) | self-reported· optimizedT1 | 2023-02-24Qwen Technical Report |
| PIQA | 65.6accuracy (%) | self-reported· optimizedT1 | 2023-02-24Qwen Technical Report |
| ARC AI2 | 59.33accuracy (%) | self-reported· optimizedT1 | 2023-02-24Qwen Technical Report |
| GSM8K | 54.4accuracy (%) | unverified· optimizedT1 | 2023-02-24Stanford HELM |
| Winogrande | 54accuracy (%) | unverified· optimizedT1 | 2023-02-24LLaMA: Open and Efficient Foundation Language Models |
| MMLU | 51.2accuracy (%) | unverified· optimizedT1 | 2023-02-24Llama 2: Open Foundation and Fine-Tuned Chat Models |
| OpenBookQA | 46.93accuracy (%) | unverified· optimizedT1 | 2023-02-24Llama 2: Open Foundation and Fine-Tuned Chat Models |
| BBH | 44.53accuracy (%) | self-reported· optimizedT1 | 2023-02-24Qwen Technical Report |
Claims drawn from cited facts, not live model generation.
This model originates from Meta AI and was released in 2023.
2 cited facts
Averaging 23.95 points behind the best disclosed scores places it clearly off the front — a blended, harness-scoped distance that bounds the story rather than telling it. Holding no top scores establishes position, not distance — a 'well behind' verdict would need gap sizes this count does not carry, since a string of near misses produces the same result.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
LLaMA-65B is an AI model developed by Meta AI, released 2023-02-24. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
LLaMA-65B has recorded scores on 10 benchmarks, each shown with its evidence status.
LLaMA-65B has recorded scores on 10 benchmarks — PIQA, OpenBookQA, Winogrande, TriviaQA, MMLU, ARC AI2, and 4 more. The full table above shows each score with its evidence status.
0 of 10 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
LLaMA-65B has 10 tracked claims: 4 self-reported, 6 unverified.