OpenAI·released 2023-06-131 source
GPT-4 (Jun 2023) benchmark scores: 10 benchmarks tracked. 40% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 10 benchmarks·best result 89.99 on GSM8K (unverified)·4 of 10 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| GSM8K | 89.99accuracy (%) | unverified· optimizedT1 | 2023-06-13Baichuan 2: Open Large-scale Language Models |
| TriviaQA | 84.8accuracy (%) | unverified· optimizedT1 | 2023-06-13PapersWithCode |
| MMLU | 76.53accuracy (%) | unverified· optimizedT1 | 2023-06-13Stanford CRFM Leaderboard |
| BBH | 66.83accuracy (%) | unverified· optimizedT1 | 2023-06-13Baichuan 2: Open Large-scale Language Models |
| METR Time Horizons | 29.3accuracy (%) | unverifiedT2 | 2023-06-13 |
| MATH level 5 | 22.97accuracy (%) | reproduced· optimizedT1 | 2023-06-13Epoch AI |
| WeirdML | 12.44accuracy (%) | unverifiedT1 | 2023-06-13WeirdML Leaderboard |
| GPQA diamond | 7.53accuracy (%) | reproduced· optimizedT1 | 2023-06-13Epoch AI |
| OTIS Mock AIME 2024-2025 | 1.01accuracy (%) | reproduced· optimizedT1 | 2023-06-13Epoch AI |
| Chess Puzzles | 0accuracy (%) | reproducedT1 | 2023-06-13Epoch AI |
Claims drawn from cited facts, not live model generation.
This model originates from OpenAI, a leading AI research organization. It was released in 2023.
2 cited facts
This model is evaluated across 10 tracked benchmarks, and it currently holds no top score on any of them. Because the scores have been independently reproduced by outside evaluation, the record can be treated as credible rather than merely vendor-claimed.
3 cited facts
With an average gap of 48.21 points behind the leader on its benchmarks, the model sits in a wide-gap band, well behind the front of the pack. It holds the top score on none of the benchmarks, indicating no category leadership on the tracked leaderboards.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
GPT-4 (Jun 2023) is an AI model developed by OpenAI, released 2023-06-13. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-4 (Jun 2023) has recorded scores on 10 benchmarks, each shown with its evidence status.
GPT-4 (Jun 2023) has recorded scores on 10 benchmarks — TriviaQA, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, MMLU, and 4 more. The full table above shows each score with its evidence status.
4 of 10 recorded scores (40%) are independently reproduced rather than self-reported by the lab.
GPT-4 (Jun 2023) has 10 tracked claims: 4 independently reproduced, 6 unverified.