OpenAI·released 2025-12-111 source
GPT-5.2 benchmark scores: 25 benchmarks tracked, leading 1. 36% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 1 of 25 benchmarks·best result 96.11 on OTIS Mock AIME 2024-2025 (reproduced)·9 of 25 independently reproduced·$1.75/$14 per M tokens
Head-to-headGPT-5.2 vs Claude Opus 4.5Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
Leads on: GDPval
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 96.11accuracy (%) | reproduced· optimizedT1 | 2025-12-11Epoch AI |
| GPQA diamond | 88.53accuracy (%) | reproduced· optimizedT1 | 2025-12-11Epoch AI |
| ARC-AGI | 86.2accuracy (%) | unverified· optimizedT1 | 2025-12-11ARC Prize Leaderboard |
| VPCT | 76accuracy (%) | unverifiedT1 | 2025-12-11VPCT leaderboard |
| METR Time Horizons | 75.29accuracy (%) | unverifiedT2 | 2025-12-11 |
| SWE-Bench verified | 73.76accuracy (%) | reproduced· optimizedT1 | 2025-12-11Epoch AI |
| WeirdML | 72.2accuracy (%) | unverifiedT1 | 2025-12-11WeirdML Leaderboard |
| FrontierMath-Tiers-1-3-v2-Private | 67.4accuracy (%) | reproducedT1 | 2025-12-11Epoch AI |
| Terminal Bench | 64.9accuracy (%) | unverified· optimizedT1 | 2025-12-11Terminal-Bench v2 Leaderboard |
| ARC-AGI-2 | 52.91accuracy (%) | unverifiedT2 | 2025-12-11 |
| GDPval | 49.7accuracy (%) | unverifiedT2 | 2025-12-11 |
| Chess Puzzles | 46.34accuracy (%) | reproducedT1 | 2025-12-11Epoch AI |
| DeepResearch Bench | 41.12accuracy (%) | unverifiedT1 | 2025-12-11https://drb.futuresearch.ai/#drb |
| SimpleQA Verified | 38.9accuracy (%) | reproducedT1 | 2025-12-11Epoch AI |
| SimpleBench | 34.96accuracy (%) | unverifiedT1 | 2025-12-11SimpleBench Leaderboard |
| APEX-Agents | 34.4accuracy (%) | unverifiedT2 | 2025-12-11 |
| FrontierMath-Tier-4-v2-Private | 31.7accuracy (%) | reproducedT1 | 2025-12-11Epoch AI |
| GSO-Bench | 27.4accuracy (%) | unverifiedT1 | 2025-12-11GSO Leaderboard |
| HLE | 24.16accuracy (%) | unverifiedT2 | 2025-12-11 |
| EBR-bench | 23.02accuracy (%) | reproducedT1 | 2025-12-11Epoch AI |
| PostTrainBench | 21.38accuracy (%) | unverifiedT2 | 2025-12-11 |
| CL-bench | 18.2accuracy (%) | unverifiedT2 | 2025-12-11 |
| Mystery Game Puzzles | 15.18accuracy (%) | reproducedT1 | 2025-12-11Epoch AI |
| ProofBench | 15accuracy (%) | unverifiedT1 | 2025-12-11https://www.vals.ai/benchmarks/proof_bench |
| Remote Labor Index | 2.5accuracy (%) | unverifiedT2 | 2025-12-11 |
Claims drawn from cited facts, not live model generation.
The model is scored on 25 tracked benchmarks. It currently holds the top score on 1 of those benchmarks. Because the record is independently reproduced, outside evaluation confirms the scores, so the numbers can be treated as verified.
3 cited facts
Relative to the best scores on the benchmarks, this model trails the leader by an average of 22.29 points under the disclosed harness — a wide gap that places it well behind the front of the pack. It still tops the leaderboard on 1 benchmark, so the gap reflects an overall deficit rather than an absence of any leading result.
2 cited facts
6 facts cross-checked across data sources: 2 corroborated, 2 single-source, 2 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, modelsdev_models, openrouter_models
GPT-5.2 is an AI model developed by OpenAI, released 2025-12-11. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-5.2 has recorded scores on 25 benchmarks, each shown with its evidence status.
GPT-5.2 has recorded scores on 25 benchmarks — GPQA diamond, OTIS Mock AIME 2024-2025, WeirdML, Chess Puzzles, SimpleBench, METR Time Horizons, and 19 more. The full table above shows each score with its evidence status.
9 of 25 recorded scores (36%) are independently reproduced rather than self-reported by the lab.
GPT-5.2 has 25 tracked claims: 9 independently reproduced, 16 unverified.
Listed API pricing: $1.75 per million input tokens, $14 per million output tokens (prices disputed across sources). See the pricing block for the full breakdown.