Meta AI·released 2024-09-241 source
Llama 3.2 90B benchmark scores: 6 benchmarks tracked. 50% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 6 benchmarks·best result 73.73 on MMLU (unverified)·3 of 6 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MMLU | 73.73accuracy (%) | unverified· optimizedT1 | 2024-09-24Stanford CRFM Leaderboard |
| GeoBench | 52accuracy (%) | unverifiedT1 | 2024-09-24GeoBench leaderboard |
| MATH level 5 | 39.44accuracy (%) | reproduced· optimizedT1 | 2024-09-24Epoch AI |
| Balrog | 27.3accuracy (%) | unverifiedT1 | 2024-09-24Balrog Leaderboard |
| GPQA diamond | 21.38accuracy (%) | reproduced· optimizedT1 | 2024-09-24Epoch AI |
| OTIS Mock AIME 2024-2025 | 2.54accuracy (%) | reproduced· optimizedT1 | 2024-09-24Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by Meta AI and released in 2024.
2 cited facts
Sparse evidence — 6 benchmarks scored, no lead — supports orientation, not judgment. Because its scores are independently reproduced, the record can be trusted.
3 cited facts
The model trails the state-of-the-art leader by an average of 50.79 points, a wide gap that places it well behind the frontier. It holds the top score on none of its benchmarks, consistent with being far from category leadership.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Llama 3.2 90B is an AI model developed by Meta AI, released 2024-09-24. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Llama 3.2 90B has recorded scores on 6 benchmarks, each shown with its evidence status.
Llama 3.2 90B has recorded scores on 6 benchmarks — GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, MMLU, Balrog, GeoBench. The full table above shows each score with its evidence status.
3 of 6 recorded scores (50%) are independently reproduced rather than self-reported by the lab.
Llama 3.2 90B has 6 tracked claims: 3 independently reproduced, 3 unverified.