OpenAI·released 2023-11-061 source
GPT-4 Turbo (Nov 2023) benchmark scores: 4 benchmarks tracked. 50% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 4 benchmarks·best result 72.8 on MMLU (unverified)·2 of 4 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MMLU | 72.8accuracy (%) | unverified· optimizedT1 | 2023-11-06Stanford CRFM Leaderboard |
| METR Time Horizons | 40.43accuracy (%) | unverifiedT1 | 2023-11-06METR - Measuring AI Ability to Complete Long Tasks |
| MATH level 5 | 40.02accuracy (%) | reproduced· optimizedT1 | 2023-11-06Epoch AI |
| GPQA diamond | 23.15accuracy (%) | reproduced· optimizedT1 | 2023-11-06Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by OpenAI and released in 2023.
2 cited facts
This model is scored on four tracked benchmarks. It holds no current top score on any of them. Because the scores are independently reproduced, the record can be trusted as verified.
3 cited facts
This model trails the current best by 44.45 points on average, a wide gap that places it well behind the leaders rather than at the frontier. It holds the top score on no benchmarks (0 leaders), so it has no category leadership yet.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
GPT-4 Turbo (Nov 2023) is an AI model developed by OpenAI, released 2023-11-06. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-4 Turbo (Nov 2023) has recorded scores on 4 benchmarks, each shown with its evidence status.
GPT-4 Turbo (Nov 2023) has recorded scores on 4 benchmarks — GPQA diamond, MATH level 5, MMLU, METR Time Horizons. The full table above shows each score with its evidence status.
2 of 4 recorded scores (50%) are independently reproduced rather than self-reported by the lab.
GPT-4 Turbo (Nov 2023) has 4 tracked claims: 2 independently reproduced, 2 unverified.