Meta AI·released 2025-04-051 source
Llama 4 Maverick benchmark scores: 14 benchmarks tracked. 29% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 14 benchmarks·best result 73.02 on MATH level 5 (reproduced)·4 of 14 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MATH level 5 | 73.02accuracy (%) | reproduced· optimizedT1 | 2025-04-05Epoch AI |
| Lech Mazur Writing | 62accuracy (%) | unverifiedT1 | 2025-04-05lechmazur/writing Github repository |
| GPQA diamond | 55.98accuracy (%) | reproduced· optimizedT1 | 2025-04-05Epoch AI |
| GeoBench | 52accuracy (%) | unverifiedT1 | 2025-04-05GeoBench leaderboard |
| Fiction.LiveBench | 46.2accuracy (%) | unverifiedT1 | 2025-04-05Fiction.live leaderboard |
| WeirdML | 24.47accuracy (%) | unverifiedT1 | 2025-04-05WeirdML Leaderboard |
| OTIS Mock AIME 2024-2025 | 20.48accuracy (%) | reproduced· optimizedT1 | 2025-04-05Epoch AI |
| Aider polyglot | 15.6accuracy (%) | unverified· optimizedT1 | 2025-04-05Aider LLM Leaderboards |
| SimpleBench | 13.24accuracy (%) | unverifiedT1 | 2025-04-05SimpleBench Leaderboard |
| ARC-AGI | 4.4accuracy (%) | unverified· optimizedT1 | 2025-04-05ARC Prize Leaderboard |
| FrontierMath-2025-02-28-Private | 1.21accuracy (%) | reproduced· optimizedT1 | 2025-04-05Epoch AI |
| HLE | 0.92accuracy (%) | unverifiedT2 | 2025-04-05 |
| CritPt | 0accuracy (%) | unverifiedT2 | 2025-04-05 |
| ARC-AGI-2 | 0accuracy (%) | unverifiedT2 | 2025-04-05 |
Claims drawn from cited facts, not live model generation.
This model was developed by Meta AI. It was released in 2025.
2 cited facts
Solid coverage at 14 tracked benchmarks makes the no-lead reading dependable as far as it goes — which is only to the top of each table. Because the record is independently reproduced, outside evaluation backs the scores up.
3 cited facts
The model trails the state of the art by an average of 55.9 points across benchmarks, placing it well behind the leaders, and it holds no category-leading scores.
2 cited facts
5 facts cross-checked across data sources: 5 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, openrouter_models
Llama 4 Maverick is an AI model developed by Meta AI, released 2025-04-05. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Llama 4 Maverick has recorded scores on 14 benchmarks, each shown with its evidence status.
Llama 4 Maverick has recorded scores on 14 benchmarks — Lech Mazur Writing, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, SimpleBench, and 8 more. The full table above shows each score with its evidence status.
4 of 14 recorded scores (29%) are independently reproduced rather than self-reported by the lab.
Llama 4 Maverick has 14 tracked claims: 4 independently reproduced, 10 unverified.