OpenAI·released 2024-05-131 source
GPT-4o (May 2024) benchmark scores: 6 benchmarks tracked, leading 1. 50% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 1 of 6 benchmarks·best result 84.67 on ScienceQA (self-reported)·3 of 6 independently reproduced
Head-to-headGPT-4o (May 2024) vs Claude 3 HaikuLeads on: ScienceQA
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| ScienceQA | 84.67accuracy (%) | self-reported· optimizedT1 | 2024-05-13Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| MMLU | 78.93accuracy (%) | unverified· optimizedT1 | 2024-05-13Stanford CRFM Leaderboard |
| MATH level 5 | 51.05accuracy (%) | reproduced· optimizedT1 | 2024-05-13Epoch AI |
| Balrog | 32.3accuracy (%) | unverifiedT1 | 2024-05-13Balrog Leaderboard |
| GPQA diamond | 31.86accuracy (%) | reproduced· optimizedT1 | 2024-05-13Epoch AI |
| OTIS Mock AIME 2024-2025 | 6.16accuracy (%) | reproduced· optimizedT1 | 2024-05-13Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by OpenAI and released in 2024.
2 cited facts
Only 6 tracked scores exist for it, so weight any pattern here lightly — the base is thin. One lead is narrow leadership — check which harness produced it before reading it as a pattern. Because its scores are independently reproduced, readers can trust the record.
3 cited facts
The model trails the state-of-the-art by an average of 38.81 points, placing it well behind the leading models.
1 cited fact
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
GPT-4o (May 2024) is an AI model developed by OpenAI, released 2024-05-13. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-4o (May 2024) has recorded scores on 6 benchmarks, each shown with its evidence status.
GPT-4o (May 2024) has recorded scores on 6 benchmarks — GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, ScienceQA, MMLU, Balrog. The full table above shows each score with its evidence status.
3 of 6 recorded scores (50%) are independently reproduced rather than self-reported by the lab.
GPT-4o (May 2024) has 6 tracked claims: 3 independently reproduced, 1 self-reported, 2 unverified.