OpenAI·released 2023-06-131 source
GPT-3.5 Turbo (Nov 2023) benchmark scores: 8 benchmarks tracked, leading 1. 25% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 1 of 8 benchmarks·best result 85.8 on TriviaQA (self-reported)·2 of 8 independently reproduced
Head-to-headGPT-3.5 Turbo (Nov 2023) vs Llama 3-8BLeads on: ANLI
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| TriviaQA | 85.8accuracy (%) | self-reported· optimizedT1 | 2023-06-13Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| ARC AI2 | 83.2accuracy (%) | self-reported· optimizedT1 | 2023-06-13Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| OpenBookQA | 81.33accuracy (%) | self-reported· optimizedT1 | 2023-06-13Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| MMLU | 61.87accuracy (%) | self-reported· optimizedT1 | 2023-06-13Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| Winogrande | 37.6accuracy (%) | self-reported· optimizedT1 | 2023-06-13Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| ANLI | 37.15accuracy (%) | self-reported· optimizedT1 | 2023-06-13Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| MATH level 5 | 15.89accuracy (%) | reproduced· optimizedT1 | 2023-06-13Epoch AI |
| GPQA diamond | 4.04accuracy (%) | reproduced· optimizedT1 | 2023-06-13Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by OpenAI and released in 2023.
2 cited facts
The verification is the headline here: outside evaluators have reproduced these scores, so the 8-benchmark record and the 1 top score it holds read as measured results, not vendor claims.
3 cited facts
This model trails the state-of-the-art leader by an average of 31.13 points, placing it well behind the front-of-pack models. Despite this wide gap overall, it holds the top score on one benchmark, showing isolated strength.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
GPT-3.5 Turbo (Nov 2023) is an AI model developed by OpenAI, released 2023-06-13. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-3.5 Turbo (Nov 2023) has recorded scores on 8 benchmarks, each shown with its evidence status.
GPT-3.5 Turbo (Nov 2023) has recorded scores on 8 benchmarks — OpenBookQA, Winogrande, TriviaQA, GPQA diamond, MATH level 5, MMLU, and 2 more. The full table above shows each score with its evidence status.
2 of 8 recorded scores (25%) are independently reproduced rather than self-reported by the lab.
GPT-3.5 Turbo (Nov 2023) has 8 tracked claims: 2 independently reproduced, 6 self-reported.