OpenAI·released 2025-08-04✓ 2 sources
gpt-oss-120b benchmark scores: 14 benchmarks tracked. 29% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 14 benchmarks·best result 88.88 on OTIS Mock AIME 2024-2025 (reproduced)·4 of 14 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 88.88accuracy (%) | reproduced· optimizedT1 | 2025-08-05Epoch AI |
| Lech Mazur Writing | 77.1accuracy (%) | unverifiedT1 | 2025-08-05lechmazur/writing Github repository |
| GPQA diamond | 67.68accuracy (%) | reproduced· optimizedT1 | 2025-08-05Epoch AI |
| METR Time Horizons | 56.63accuracy (%) | unverifiedT1 | 2025-08-05METR - Measuring AI Ability to Complete Long Tasks |
| WeirdML | 48.17accuracy (%) | unverifiedT1 | 2025-08-05WeirdML Leaderboard |
| Fiction.LiveBench | 44.4accuracy (%) | unverifiedT1 | 2025-08-05Fiction.live leaderboard |
| Aider polyglot | 41.8accuracy (%) | unverified· optimizedT1 | 2025-08-05Aider LLM Leaderboards |
| Surface Evolver Bench | 25accuracy (%) | unverifiedT1 | 2025-08-05https://yhenon.github.io/surface-evolver-llm-eval/ |
| Terminal Bench | 18.7accuracy (%) | unverified· optimizedT1 | 2025-08-05Terminal-Bench v2 Leaderboard |
| Chess Puzzles | 15.82accuracy (%) | reproducedT1 | 2025-08-05Epoch AI |
| SimpleQA Verified | 13.9accuracy (%) | reproducedT1 | 2025-08-05Epoch AI |
| SimpleBench | 6.52accuracy (%) | unverifiedT1 | 2025-08-05SimpleBench Leaderboard |
| APEX-Agents | 4.7accuracy (%) | unverifiedT2 | 2025-08-05 |
| CritPt | 1.14accuracy (%) | unverifiedT2 | 2025-08-05 |
Claims drawn from cited facts, not live model generation.
Developed by OpenAI and released in 2025, this model is a frontier-scale system with 116.83 billion parameters. That scale implies ample headroom for complex tasks, but also substantial computational expense.
3 cited facts
On average this model trails the SOTA leader by 42.81 points, a wide gap indicating it is well behind the leaders on its benchmarks. It currently holds the top score on none of the tracked benchmarks, so there is no category leadership yet.
2 cited facts
6 facts cross-checked across data sources: 1 corroborated, 5 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, huggingface_models, openrouter_models
gpt-oss-120b is an AI model developed by OpenAI, released 2025-08-04. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
gpt-oss-120b has recorded scores on 14 benchmarks, each shown with its evidence status.
gpt-oss-120b has recorded scores on 14 benchmarks — Lech Mazur Writing, GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleBench, and 8 more. The full table above shows each score with its evidence status.
4 of 14 recorded scores (29%) are independently reproduced rather than self-reported by the lab.
gpt-oss-120b has 14 tracked claims: 4 independently reproduced, 10 unverified.