OpenAI·released 2023-06-131 source
GPT-3.5 Turbo (Jun 2023) benchmark scores: 4 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 4 benchmarks·best result 63.2 on Winogrande (unverified)·0 of 4 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| Winogrande | 63.2accuracy (%) | unverified· optimizedT1 | 2023-06-13The Falcon Series of Open Language Models |
| MMLU | 58.53accuracy (%) | unverified· optimizedT1 | 2023-06-13Stanford CRFM Leaderboard |
| GSM8K | 57.77accuracy (%) | unverified· optimizedT1 | 2023-06-13Stanford HELM |
| BBH | 48.79accuracy (%) | unverified· optimizedT1 | 2023-06-13Baichuan 2: Open Large-scale Language Models |
Claims drawn from cited facts, not live model generation.
This model was developed by OpenAI and first released in 2023.
2 cited facts
This model is scored on 4 tracked benchmarks and holds no current top score; its record is unverified, meaning it is neither confirmed nor disputed by outside evaluation.
3 cited facts
This model trails the state-of-the-art by an average of 27.96 points, which places it well behind the leaders.
1 cited fact
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
GPT-3.5 Turbo (Jun 2023) is an AI model developed by OpenAI, released 2023-06-13. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-3.5 Turbo (Jun 2023) has recorded scores on 4 benchmarks, each shown with its evidence status.
GPT-3.5 Turbo (Jun 2023) has recorded scores on 4 benchmarks — Winogrande, MMLU, BBH, GSM8K. The full table above shows each score with its evidence status.
0 of 4 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
GPT-3.5 Turbo (Jun 2023) has 4 tracked claims: 4 unverified.