Meta AI·released 2024-11-26✓ 2 sources
Llama 3.3 70B benchmark scores: 9 benchmarks tracked. 33% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 9 benchmarks·best result 81.73 on MMLU (unverified)·3 of 9 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MMLU | 81.73accuracy (%) | unverified· optimizedT1 | 2024-12-06Stanford CRFM Leaderboard |
| MATH level 5 | 41.6accuracy (%) | reproduced· optimizedT1 | 2024-12-06Epoch AI |
| Fiction.LiveBench | 33.3accuracy (%) | unverifiedT1 | 2024-12-06Fiction.live leaderboard |
| GPQA diamond | 29.92accuracy (%) | reproduced· optimizedT1 | 2024-12-06Epoch AI |
| Balrog | 23accuracy (%) | unverifiedT1 | 2024-12-06Balrog Leaderboard |
| WeirdML | 14.44accuracy (%) | unverifiedT1 | 2024-12-06WeirdML Leaderboard |
| OTIS Mock AIME 2024-2025 | 5.04accuracy (%) | reproduced· optimizedT1 | 2024-12-06Epoch AI |
| SimpleBench | 3.88accuracy (%) | unverifiedT1 | 2024-12-06SimpleBench Leaderboard |
| CritPt | 0accuracy (%) | unverifiedT2 | 2024-12-06 |
Claims drawn from cited facts, not live model generation.
This model was developed by Meta AI and released in 2024. With roughly 70.55 billion parameters, it is a large-scale model whose size provides greater headroom but also brings higher computational cost.
3 cited facts
This model is tracked across 9 benchmarks and currently holds no top score on any of them. Because these results have been independently reproduced, the record can be trusted as verified rather than vendor claims.
3 cited facts
Relative to the disclosed benchmark harness, the model trails the leader by an average of 55.59 points — a wide gap that places it well behind the front of the pack. It holds the top score on none of the tracked benchmarks, so there is no benchmark category where it currently claims state-of-the-art leadership.
2 cited facts
6 facts cross-checked across data sources: 1 corroborated, 5 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, huggingface_models, openrouter_models
Llama 3.3 70B is an AI model developed by Meta AI, released 2024-11-26. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Llama 3.3 70B has recorded scores on 9 benchmarks, each shown with its evidence status.
Llama 3.3 70B has recorded scores on 9 benchmarks — GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, MMLU, SimpleBench, and 3 more. The full table above shows each score with its evidence status.
3 of 9 recorded scores (33%) are independently reproduced rather than self-reported by the lab.
Llama 3.3 70B has 9 tracked claims: 3 independently reproduced, 6 unverified.