OpenAI·released 2024-05-131 source
GPT-4o (Aug 2024) benchmark scores: 10 benchmarks tracked. 50% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 10 benchmarks·best result 79.07 on MMLU (unverified)·5 of 10 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MMLU | 79.07accuracy (%) | unverified· optimizedT1 | 2024-05-13Stanford CRFM Leaderboard |
| MATH level 5 | 53.28accuracy (%) | reproduced· optimizedT1 | 2024-05-13Epoch AI |
| METR Time Horizons | 33.84accuracy (%) | unverifiedT2 | 2024-05-13 |
| GPQA diamond | 32.28accuracy (%) | reproduced· optimizedT1 | 2024-05-13Epoch AI |
| CadEval | 26accuracy (%) | unverifiedT1 | 2024-05-13CadEval Dashboard |
| Aider polyglot | 23.1accuracy (%) | unverified· optimizedT1 | 2024-05-13Aider LLM Leaderboards |
| Chess Puzzles | 8.46accuracy (%) | reproducedT1 | 2024-05-13Epoch AI |
| OTIS Mock AIME 2024-2025 | 6.3accuracy (%) | reproduced· optimizedT1 | 2024-05-13Epoch AI |
| SimpleBench | 1.36accuracy (%) | unverifiedT1 | 2024-05-13SimpleBench Leaderboard |
| FrontierMath-2025-02-28-Private | 0.6accuracy (%) | reproduced· optimizedT1 | 2024-05-13Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by OpenAI and released in 2024.
2 cited facts
It is scored on 10 tracked benchmarks and currently holds no top score on any of them. Because the submission has been independently reproduced, the public record can be treated as reliable.
3 cited facts
With an average sota gap of 56.08 points, this model sits well behind the leaders on its benchmarks, a wide gap rather than a competitive one. It holds the top score on none of the tracked benchmarks, so it shows no category leadership.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
GPT-4o (Aug 2024) is an AI model developed by OpenAI, released 2024-05-13. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-4o (Aug 2024) has recorded scores on 10 benchmarks, each shown with its evidence status.
GPT-4o (Aug 2024) has recorded scores on 10 benchmarks — GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, CadEval, MMLU, Chess Puzzles, and 4 more. The full table above shows each score with its evidence status.
5 of 10 recorded scores (50%) are independently reproduced rather than self-reported by the lab.
GPT-4o (Aug 2024) has 10 tracked claims: 5 independently reproduced, 5 unverified.