OpenAI·released 2024-05-131 source
GPT-4o (Nov 2024) benchmark scores: 21 benchmarks tracked, leading 1. 24% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 1 of 21 benchmarks·best result 84.13 on MMLU (self-reported)·5 of 21 independently reproduced
Head-to-headGPT-4o (Nov 2024) vs Claude 3.5 Sonnet (October 2024)Leads on: MMLU
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MMLU | 84.13accuracy (%) | self-reported· optimizedT1 | 2024-05-13Phi-4 Technical Report |
| Lech Mazur Writing | 81.8accuracy (%) | unverifiedT1 | 2024-05-13lechmazur/writing Github repository |
| GeoBench | 71accuracy (%) | unverifiedT1 | 2024-05-13GeoBench leaderboard |
| MATH level 5 | 49.77accuracy (%) | reproduced· optimizedT1 | 2024-05-13Epoch AI |
| METR Time Horizons | 40.76accuracy (%) | unverifiedT1 | 2024-05-13METR - Measuring AI Ability to Complete Long Tasks |
| SWE-Bench verified | 30.99accuracy (%) | reproduced· optimizedT1 | 2024-05-13Epoch AI |
| GPQA diamond | 30.51accuracy (%) | reproduced· optimizedT1 | 2024-05-13Epoch AI |
| WeirdML | 25.12accuracy (%) | unverifiedT1 | 2024-05-13WeirdML Leaderboard |
| Aider polyglot | 18.2accuracy (%) | unverified· optimizedT1 | 2024-05-13Aider LLM Leaderboards |
| Cybench | 12.5accuracy (%) | unverified· optimizedT1 | 2024-05-13Cybench leaderboard |
| VPCT | 10accuracy (%) | unverifiedT1 | 2024-05-13VPCT leaderboard |
| GDPval | 9.9accuracy (%) | unverifiedT2 | 2024-05-13 |
| The Agent Company | 8.6accuracy (%) | unverifiedT1 | 2024-05-13TheAgentCompany experiment results github |
| OTIS Mock AIME 2024-2025 | 6.16accuracy (%) | reproduced· optimizedT1 | 2024-05-13Epoch AI |
| ARC-AGI | 4.5accuracy (%) | unverified· optimizedT1 | 2024-05-13ARC Prize Leaderboard |
| APEX-Agents | 1.1accuracy (%) | unverifiedT2 | 2024-05-13 |
| FrontierMath-2025-02-28-Private | 0.6accuracy (%) | reproduced· optimizedT1 | 2024-05-13Epoch AI |
| CritPt | 0accuracy (%) | unverifiedT2 | 2024-05-13 |
| GSO-Bench | 0accuracy (%) | unverifiedT1 | 2024-05-13GSO Leaderboard |
| ARC-AGI-2 | 0accuracy (%) | unverifiedT2 | 2024-05-13 |
| HLE | 0accuracy (%) | unverifiedT2 | 2024-05-13 |
Claims drawn from cited facts, not live model generation.
This model was developed by OpenAI and released in 2024.
2 cited facts
This model is scored on 21 tracked benchmarks, and it currently holds the top score on 1 of them. Because the results have been independently reproduced, these scores can be treated as verified rather than merely vendor-claimed.
3 cited facts
Based on an average gap of 52.65 points under the disclosed harness, this model trails the leader by a wide margin in aggregate, placing it well behind the front of the pack. It does, however, hold the top score on 1 benchmark, which indicates genuine category leadership in that specific instance despite the overall gap.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
GPT-4o (Nov 2024) is an AI model developed by OpenAI, released 2024-05-13. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-4o (Nov 2024) has recorded scores on 21 benchmarks, each shown with its evidence status.
GPT-4o (Nov 2024) has recorded scores on 21 benchmarks — Lech Mazur Writing, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, MMLU, and 15 more. The full table above shows each score with its evidence status.
5 of 21 recorded scores (24%) are independently reproduced rather than self-reported by the lab.
GPT-4o (Nov 2024) has 21 tracked claims: 5 independently reproduced, 1 self-reported, 15 unverified.