OpenAI·released 2023-03-151 source
GPT-4 (Mar 2023) benchmark scores: 7 benchmarks tracked, leading 2. 29% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 2 of 7 benchmarks·best result 93.73 on HellaSwag (unverified)·2 of 7 independently reproduced
Head-to-headGPT-4 (Mar 2023) vs GPT-4o mini| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| HellaSwag | 93.73accuracy (%) | unverified· optimizedT1 | 2023-03-15The Falcon Series of Open Language Models |
| GSM8K | 92accuracy (%) | self-reported· optimizedT1 | 2023-03-15GPT-4 Technical Report |
| MMLU | 81.87accuracy (%) | self-reported· optimizedT1 | 2023-03-15GPT-4 Technical Report |
| Winogrande | 75accuracy (%) | self-reported· optimizedT1 | 2023-03-15GPT-4 technical report |
| METR Time Horizons | 36.11accuracy (%) | unverifiedT1 | 2023-03-15METR - Measuring AI Ability to Complete Long Tasks |
| GPQA diamond | 14.31accuracy (%) | reproduced· optimizedT1 | 2023-03-15Epoch AI |
| OTIS Mock AIME 2024-2025 | 0.46accuracy (%) | reproduced· optimizedT1 | 2023-03-15Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by OpenAI. It was released in 2023, and no parameter count is available, so its scale cannot be characterized beyond this origin and vintage.
2 cited facts
This model is tracked across seven benchmarks and currently holds the top score on two of them. These results have been independently reproduced, so the record can be treated as verified rather than vendor-reported.
3 cited facts
On average, this model trails the state-of-the-art leader by 32.39 points under the disclosed harness, a wide gap that places it well behind the front-of-pack despite holding the top score on 2 benchmarks.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
GPT-4 (Mar 2023) is an AI model developed by OpenAI, released 2023-03-15. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-4 (Mar 2023) has recorded scores on 7 benchmarks, each shown with its evidence status.
GPT-4 (Mar 2023) has recorded scores on 7 benchmarks — Winogrande, GPQA diamond, OTIS Mock AIME 2024-2025, MMLU, METR Time Horizons, GSM8K, and 1 more. The full table above shows each score with its evidence status.
2 of 7 recorded scores (29%) are independently reproduced rather than self-reported by the lab.
GPT-4 (Mar 2023) has 7 tracked claims: 2 independently reproduced, 3 self-reported, 2 unverified.