OpenAI·released 2023-06-131 source
GPT-3.5 Turbo (Jan 2024) benchmark scores: 6 benchmarks tracked. 67% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 6 benchmarks·best result 56.4 on MMLU (unverified)·4 of 6 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MMLU | 56.4accuracy (%) | unverified· optimizedT1 | 2023-06-13Stanford CRFM Leaderboard |
| MATH level 5 | 11.63accuracy (%) | reproduced· optimizedT1 | 2023-06-13Epoch AI |
| WeirdML | 3.48accuracy (%) | unverifiedT1 | 2023-06-13WeirdML Leaderboard |
| GPQA diamond | 2.9accuracy (%) | reproduced· optimizedT1 | 2023-06-13Epoch AI |
| OTIS Mock AIME 2024-2025 | 2.12accuracy (%) | reproduced· optimizedT1 | 2023-06-13Epoch AI |
| Chess Puzzles | 0accuracy (%) | reproducedT1 | 2023-06-13Epoch AI |
Claims drawn from cited facts, not live model generation.
This model originates from OpenAI and was released in 2023.
2 cited facts
This model is tracked on six benchmarks, and currently holds no top score in any of them. The scores are backed by independent outside evaluation, so the record can be trusted.
3 cited facts
With an average SOTA gap of 75.48 points, the model trails the leader by a wide margin on its benchmarks, placing it well behind the front of the pack. It holds the top score on zero benchmarks, so it has no category leadership yet.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
GPT-3.5 Turbo (Jan 2024) is an AI model developed by OpenAI, released 2023-06-13. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-3.5 Turbo (Jan 2024) has recorded scores on 6 benchmarks, each shown with its evidence status.
GPT-3.5 Turbo (Jan 2024) has recorded scores on 6 benchmarks — GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, MMLU, Chess Puzzles. The full table above shows each score with its evidence status.
4 of 6 recorded scores (67%) are independently reproduced rather than self-reported by the lab.
GPT-3.5 Turbo (Jan 2024) has 6 tracked claims: 4 independently reproduced, 2 unverified.