OpenAI·released 2025-04-141 source
GPT-4.1 mini benchmark scores: 12 benchmarks tracked. 42% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 12 benchmarks·best result 87.29 on MATH level 5 (reproduced)·5 of 12 independently reproduced·$0.4/$1.6 per M tokens
Source: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MATH level 5 | 87.29accuracy (%) | reproduced· optimizedT1 | 2025-04-14Epoch AI |
| GPQA diamond | 54.46accuracy (%) | reproduced· optimizedT1 | 2025-04-14Epoch AI |
| OTIS Mock AIME 2024-2025 | 44.67accuracy (%) | reproduced· optimizedT1 | 2025-04-14Epoch AI |
| Fiction.LiveBench | 44.4accuracy (%) | unverifiedT1 | 2025-04-14Fiction.live leaderboard |
| WeirdML | 37.61accuracy (%) | unverifiedT1 | 2025-04-14WeirdML Leaderboard |
| Aider polyglot | 32.4accuracy (%) | unverified· optimizedT1 | 2025-04-14Aider LLM Leaderboards |
| CadEval | 16accuracy (%) | unverifiedT1 | 2025-04-14CadEval Dashboard |
| FrontierMath-2025-02-28-Private | 7.86accuracy (%) | reproduced· optimizedT1 | 2025-04-14Epoch AI |
| ARC-AGI | 3.5accuracy (%) | unverified· optimizedT1 | 2025-04-14ARC Prize Leaderboard |
| Chess Puzzles | 2.15accuracy (%) | reproducedT1 | 2025-04-14Epoch AI |
| CritPt | 0accuracy (%) | unverifiedT2 | 2025-04-14 |
| ARC-AGI-2 | 0accuracy (%) | unverifiedT2 | 2025-04-14 |
Claims drawn from cited facts, not live model generation.
This model originates from OpenAI and was released in 2025. With no parameter count disclosed, its scale cannot be characterized, but its vintage places it among the developer's current generation.
3 cited facts
This model is scored on 12 tracked benchmarks. Among those, it holds no current top score. The record is independently reproduced, so outside evaluation backs the scores up.
3 cited facts
Based on the available derived metric, the model trails the top score by an average of 55.49 points, a wide gap that places it well behind the leaders rather than at the frontier. It does not hold the top score on any of the tracked benchmarks, meaning it has no demonstrated category leadership among those leaders.
2 cited facts
6 facts cross-checked across data sources: 3 corroborated, 2 single-source, 1 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, modelsdev_models, openrouter_models
GPT-4.1 mini is an AI model developed by OpenAI, released 2025-04-14. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-4.1 mini has recorded scores on 12 benchmarks, each shown with its evidence status.
GPT-4.1 mini has recorded scores on 12 benchmarks — GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, CadEval, Chess Puzzles, and 6 more. The full table above shows each score with its evidence status.
5 of 12 recorded scores (42%) are independently reproduced rather than self-reported by the lab.
GPT-4.1 mini has 12 tracked claims: 5 independently reproduced, 7 unverified.
Listed API pricing: $0.4 per million input tokens, $1.6 per million output tokens. See the pricing block for the full breakdown.