OpenAI·released 2025-08-04✓ 2 sources
gpt-oss-20b benchmark scores: 5 benchmarks tracked. 40% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 5 benchmarks·best result 53.84 on OTIS Mock AIME 2024-2025 (reproduced)·2 of 5 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 53.84accuracy (%) | reproduced· optimizedT1 | 2025-08-05Epoch AI |
| WeirdML | 40.93accuracy (%) | unverifiedT1 | 2025-08-05WeirdML Leaderboard |
| GPQA diamond | 27.95accuracy (%) | reproduced· optimizedT1 | 2025-08-05Epoch AI |
| Terminal Bench | 3.4accuracy (%) | unverified· optimizedT1 | 2025-08-05Terminal-Bench v2 Leaderboard |
| CritPt | 1.43accuracy (%) | unverifiedT2 | 2025-08-05 |
Claims drawn from cited facts, not live model generation.
Developed by OpenAI and released in 2025, this model is a mid-size system with 21.51 billion parameters. That scale implies a practical balance: it offers more headroom than compact models while remaining more efficient and less expensive to run than frontier-scale alternatives.
4 cited facts
This model is scored on 5 tracked benchmarks. It holds no current top score on any of those benchmarks. Because these results have been independently reproduced, the record can be trusted as verified.
3 cited facts
On average, this model trails the state-of-the-art leader by 54.9 points under the disclosed evaluation harness, a wide gap that places it well behind front-of-pack systems. It holds the top score on none of the sota leader count benchmarks, so there is no benchmark category where it currently stands as leader.
2 cited facts
2 facts cross-checked across data sources: 1 corroborated, 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, huggingface_models
gpt-oss-20b is an AI model developed by OpenAI, released 2025-08-04. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
gpt-oss-20b has recorded scores on 5 benchmarks, each shown with its evidence status.
gpt-oss-20b has recorded scores on 5 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, CritPt, Terminal Bench. The full table above shows each score with its evidence status.
2 of 5 recorded scores (40%) are independently reproduced rather than self-reported by the lab.
gpt-oss-20b has 5 tracked claims: 2 independently reproduced, 3 unverified.