OpenAI·released 2026-02-051 source
GPT-5.3 Codex benchmark scores: 6 benchmarks tracked. 17% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 6 benchmarks·best result 79.3 on WeirdML (unverified)·1 of 6 independently reproduced·$1.75/$14 per M tokens
Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| WeirdML | 79.3accuracy (%) | unverifiedT1 | 2026-02-05WeirdML Leaderboard |
| Terminal Bench | 78.4accuracy (%) | unverified· optimizedT1 | 2026-02-05https://www.tbench.ai/leaderboard/terminal-bench/2.0 |
| SWE-Bench verified | 74.79accuracy (%) | reproduced· optimizedT1 | 2026-02-05Epoch AI |
| METR Time Horizons | 74.54accuracy (%) | unverifiedT2 | 2026-02-05 |
| APEX-Agents | 31.8accuracy (%) | unverifiedT2 | 2026-02-05 |
| PostTrainBench | 17.76accuracy (%) | unverifiedT2 | 2026-02-05 |
Claims drawn from cited facts, not live model generation.
OpenAI developed this model, placing its origin with a leading AI research lab. Released in 2026, this model's vintage reflects a current-generation release.
2 cited facts
This model is evaluated on six tracked benchmarks. It holds no current top score on any of them. The record is independently reproduced, so outside evaluation backs these numbers and they can be read as verified.
3 cited facts
With an average gap of 11.53 points behind the leader, this model sits well behind the front of the pack on its benchmarks. It holds the top score on none of the tracked benchmarks, so it does not currently claim category leadership.
2 cited facts
6 facts cross-checked across data sources: 2 corroborated, 2 single-source, 2 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, modelsdev_models, openrouter_models
GPT-5.3 Codex is an AI model developed by OpenAI, released 2026-02-05. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-5.3 Codex has recorded scores on 6 benchmarks, each shown with its evidence status.
GPT-5.3 Codex has recorded scores on 6 benchmarks — WeirdML, METR Time Horizons, SWE-Bench verified, APEX-Agents, PostTrainBench, Terminal Bench. The full table above shows each score with its evidence status.
1 of 6 recorded scores (17%) are independently reproduced rather than self-reported by the lab.
GPT-5.3 Codex has 6 tracked claims: 1 independently reproduced, 5 unverified.
Listed API pricing: $1.75 per million input tokens, $14 per million output tokens (prices disputed across sources). See the pricing block for the full breakdown.