Meta AI·released 2023-02-271 source
LLaMA-33B benchmark scores: 10 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 10 benchmarks·best result 83.8 on TriviaQA (unverified)·0 of 10 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| TriviaQA | 83.8accuracy (%) | unverified· optimizedT1 | 2023-02-27Llama 2: Open Foundation and Fine-Tuned Chat Models |
| LAMBADA | 77.2accuracy (%) | self-reported· optimizedT1 | 2023-02-27Qwen Technical Report |
| HellaSwag | 77.07accuracy (%) | unverified· optimizedT1 | 2023-02-27LLaMA: Open and Efficient Foundation Language Models |
| PIQA | 64.6accuracy (%) | self-reported· optimizedT1 | 2023-02-27Qwen Technical Report |
| ARC AI2 | 56.67accuracy (%) | self-reported· optimizedT1 | 2023-02-27Qwen Technical Report |
| Winogrande | 52accuracy (%) | unverified· optimizedT1 | 2023-02-27LLaMA: Open and Efficient Foundation Language Models |
| MMLU | 44.93accuracy (%) | unverified· optimizedT1 | 2023-02-27Llama 2: Open Foundation and Fine-Tuned Chat Models |
| OpenBookQA | 44.8accuracy (%) | unverified· optimizedT1 | 2023-02-27Llama 2: Open Foundation and Fine-Tuned Chat Models |
| GSM8K | 44.1accuracy (%) | unverified· optimizedT1 | 2023-02-27Stanford HELM |
| BBH | 33.33accuracy (%) | self-reported· optimizedT1 | 2023-02-27Qwen Technical Report |
Claims drawn from cited facts, not live model generation.
This model was developed by Meta AI and released in 2023.
2 cited facts
Middling breadth and no front-runner cell — a 10-benchmark record like this is context for the table beside it, not a standalone verdict. Because the record is self-reported, these numbers are vendor-claimed and not independently confirmed, so treat them as claims rather than verified results.
3 cited facts
This model trails the state-of-the-art by an average of 27.96 points, a wide gap that places it well behind the leading models.
1 cited fact
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
LLaMA-33B is an AI model developed by Meta AI, released 2023-02-27. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
LLaMA-33B has recorded scores on 10 benchmarks, each shown with its evidence status.
LLaMA-33B has recorded scores on 10 benchmarks — PIQA, OpenBookQA, Winogrande, TriviaQA, MMLU, ARC AI2, and 4 more. The full table above shows each score with its evidence status.
0 of 10 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
LLaMA-33B has 10 tracked claims: 4 self-reported, 6 unverified.