OpenAI·released 2025-04-141 source
GPT-4.1 benchmark scores: 17 benchmarks tracked. 41% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 17 benchmarks·best result 83.01 on MATH level 5 (reproduced)·7 of 17 independently reproduced·$2/$8 per M tokens
Consensus: LiteLLM · Cross-check: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MATH level 5 | 83.01accuracy (%) | reproduced· optimizedT1 | 2025-04-14Epoch AI |
| GeoBench | 72accuracy (%) | unverifiedT1 | 2025-04-14GeoBench leaderboard |
| Fiction.LiveBench | 63.9accuracy (%) | unverifiedT1 | 2025-04-14Fiction.live leaderboard |
| GPQA diamond | 55.89accuracy (%) | reproduced· optimizedT1 | 2025-04-14Epoch AI |
| Aider polyglot | 52.4accuracy (%) | unverified· optimizedT1 | 2025-04-14Aider LLM Leaderboards |
| SWE-Bench verified | 48.54accuracy (%) | reproduced· optimizedT1 | 2025-04-14Epoch AI |
| CadEval | 42accuracy (%) | unverifiedT1 | 2025-04-14CadEval Dashboard |
| WeirdML | 39.04accuracy (%) | unverifiedT1 | 2025-04-14WeirdML Leaderboard |
| OTIS Mock AIME 2024-2025 | 38.27accuracy (%) | reproduced· optimizedT1 | 2025-04-14Epoch AI |
| DeepResearch Bench | 29.31accuracy (%) | unverifiedT2 | 2025-04-14 |
| SimpleBench | 12.4accuracy (%) | unverifiedT1 | 2025-04-14SimpleBench Leaderboard |
| FrontierMath-2025-02-28-Private | 9.68accuracy (%) | reproduced· optimizedT1 | 2025-04-14Epoch AI |
| ARC-AGI | 5.5accuracy (%) | unverified· optimizedT1 | 2025-04-14ARC Prize Leaderboard |
| Chess Puzzles | 1.09accuracy (%) | reproducedT1 | 2025-04-14Epoch AI |
| HLE | 0.63accuracy (%) | unverifiedT2 | 2025-04-14 |
| ARC-AGI-2 | 0.42accuracy (%) | unverifiedT2 | 2025-04-14 |
| FrontierMath-Tier-4-2025-07-01-Private | 0accuracy (%) | reproducedT1 | 2025-04-14Epoch AI |
Claims drawn from cited facts, not live model generation.
This model, developed by OpenAI, was released in 2025.
2 cited facts
This model is scored on 17 tracked benchmarks. It holds no current top score on any of them. Because the results have been independently reproduced, the record can be trusted as verified.
3 cited facts
With an average gap of 46.46 points to the leader, this model sits well behind the front of the pack. It holds the top score on none of the benchmarks, so it has no category leadership to claim.
2 cited facts
6 facts cross-checked across data sources: 3 corroborated, 1 single-source, 2 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, litellm_prices, modelsdev_models, openrouter_models
GPT-4.1 is an AI model developed by OpenAI, released 2025-04-14. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-4.1 has recorded scores on 17 benchmarks, each shown with its evidence status.
GPT-4.1 has recorded scores on 17 benchmarks — GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, CadEval, Chess Puzzles, and 11 more. The full table above shows each score with its evidence status.
7 of 17 recorded scores (41%) are independently reproduced rather than self-reported by the lab.
GPT-4.1 has 17 tracked claims: 7 independently reproduced, 10 unverified.
Listed API pricing: $2 per million input tokens, $8 per million output tokens. See the pricing block for the full breakdown.