OpenAI·released 2025-11-191 source
GPT-5.1-Codex-Max benchmark scores: 4 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 4 benchmarks·best result 71.58 on METR Time Horizons (unverified)·0 of 4 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| METR Time Horizons | 71.58accuracy (%) | unverifiedT1 | 2025-11-19METR - Measuring AI Ability to Complete Long Tasks |
| Terminal Bench | 60.4accuracy (%) | unverified· optimizedT1 | 2025-11-19Terminal-Bench v2 Leaderboard |
| PostTrainBench | 19.68accuracy (%) | unverifiedT2 | 2025-11-19 |
| ProofBench | 9accuracy (%) | unverifiedT1 | 2025-11-19https://www.vals.ai/benchmarks/proof_bench |
Claims drawn from cited facts, not live model generation.
This model originates from OpenAI's research and development program and was released in 2025.
2 cited facts
This model is tracked on 4 benchmarks and holds no current top score on any of them. Because the record is unverified, these figures should be read as provisional rather than independently confirmed.
3 cited facts
Based on the average gap to the state of the art of 35.92 points, the model trails the leader by roughly 36 points on average, which is a wide gap and places it well behind the front-of-pack on its benchmarks (under the disclosed harness). The model holds the top score on zero benchmarks, meaning it has no genuine category leadership yet.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
GPT-5.1-Codex-Max is an AI model developed by OpenAI, released 2025-11-19. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-5.1-Codex-Max has recorded scores on 4 benchmarks, each shown with its evidence status.
GPT-5.1-Codex-Max has recorded scores on 4 benchmarks — METR Time Horizons, PostTrainBench, ProofBench, Terminal Bench. The full table above shows each score with its evidence status.
0 of 4 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
GPT-5.1-Codex-Max has 4 tracked claims: 4 unverified.