Meta AI·released 2025-04-051 source
Llama 4 Scout benchmark scores: 8 benchmarks tracked. 50% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 8 benchmarks·best result 62.27 on MATH level 5 (reproduced)·4 of 8 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MATH level 5 | 62.27accuracy (%) | reproduced· optimizedT1 | 2025-04-05Epoch AI |
| Fiction.LiveBench | 36accuracy (%) | unverifiedT1 | 2025-04-05Fiction.live leaderboard |
| GPQA diamond | 35.77accuracy (%) | reproduced· optimizedT1 | 2025-04-05Epoch AI |
| OTIS Mock AIME 2024-2025 | 7.69accuracy (%) | reproduced· optimizedT1 | 2025-04-05Epoch AI |
| ARC-AGI | 0.5accuracy (%) | unverified· optimizedT1 | 2025-04-05ARC Prize Leaderboard |
| FrontierMath-2025-02-28-Private | 0accuracy (%) | reproduced· optimizedT1 | 2025-04-05Epoch AI |
| CritPt | 0accuracy (%) | unverifiedT2 | 2025-04-05 |
| ARC-AGI-2 | 0accuracy (%) | unverifiedT2 | 2025-04-05 |
Claims drawn from cited facts, not live model generation.
This model was developed by Meta AI and released in 2025.
2 cited facts
This model is scored on 8 tracked benchmarks, holds no current top score, and because the record is independently reproduced, the scores are backed by outside evaluation and can be trusted as verified.
3 cited facts
A 65.99-point average deficit puts the tracked frontier far out of reach in this scoring — read it with its denominator in mind, since which benchmarks get averaged sets how dramatic the number looks.
2 cited facts
5 facts cross-checked across data sources: 5 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, openrouter_models
Llama 4 Scout is an AI model developed by Meta AI, released 2025-04-05. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Llama 4 Scout has recorded scores on 8 benchmarks, each shown with its evidence status.
Llama 4 Scout has recorded scores on 8 benchmarks — GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, FrontierMath-2025-02-28-Private, CritPt, ARC-AGI, and 2 more. The full table above shows each score with its evidence status.
4 of 8 recorded scores (50%) are independently reproduced rather than self-reported by the lab.
Llama 4 Scout has 8 tracked claims: 4 independently reproduced, 4 unverified.