OpenAI·released 2025-02-271 source
GPT-4.5 benchmark scores: 13 benchmarks tracked. 23% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 13 benchmarks·best result 78.63 on MATH level 5 (reproduced)·3 of 13 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MATH level 5 | 78.63accuracy (%) | reproduced· optimizedT1 | 2025-02-27Epoch AI |
| Lech Mazur Writing | 75.6accuracy (%) | unverifiedT1 | 2025-02-27lechmazur/writing Github repository |
| Fiction.LiveBench | 63.9accuracy (%) | unverifiedT1 | 2025-02-27Fiction.live leaderboard |
| GPQA diamond | 58.25accuracy (%) | reproduced· optimizedT1 | 2025-02-27Epoch AI |
| Aider polyglot | 44.9accuracy (%) | unverified· optimizedT1 | 2025-02-27Aider LLM Leaderboards |
| WeirdML | 39.37accuracy (%) | unverifiedT1 | 2025-02-27WeirdML Leaderboard |
| OTIS Mock AIME 2024-2025 | 37.72accuracy (%) | reproduced· optimizedT1 | 2025-02-27Epoch AI |
| SimpleBench | 21.4accuracy (%) | unverifiedT1 | 2025-02-27SimpleBench Leaderboard |
| Cybench | 17.5accuracy (%) | unverified· optimizedT1 | 2025-02-27Cybench leaderboard |
| VPCT | 17.5accuracy (%) | unverifiedT1 | 2025-02-27VPCT leaderboard |
| ARC-AGI | 10.3accuracy (%) | unverified· optimizedT1 | 2025-02-27ARC Prize Leaderboard |
| ARC-AGI-2 | 0.8accuracy (%) | unverifiedT2 | 2025-02-27 |
| HLE | 0.67accuracy (%) | unverifiedT2 | 2025-02-27 |
Claims drawn from cited facts, not live model generation.
This model was developed by OpenAI and released in 2025.
2 cited facts
13 benchmarks is enough width that a stray anomaly should not steer the read; the cross-table pattern is the signal. Currently at the front of nothing it is tracked on — that bounds what this record can claim, and says nothing about where it sits below the top. Because the claim status is independently reproduced, independent evaluation backs up these scores, so the record is trustworthy.
3 cited facts
The model trails the state-of-the-art leader by an average of 51.48 points, a wide gap that places it well behind the front-runners. It holds the top score on zero benchmarks, indicating no category leadership.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
GPT-4.5 is an AI model developed by OpenAI, released 2025-02-27. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-4.5 has recorded scores on 13 benchmarks, each shown with its evidence status.
GPT-4.5 has recorded scores on 13 benchmarks — Lech Mazur Writing, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, Cybench, and 7 more. The full table above shows each score with its evidence status.
3 of 13 recorded scores (23%) are independently reproduced rather than self-reported by the lab.
GPT-4.5 has 13 tracked claims: 3 independently reproduced, 10 unverified.