OpenAI·released 2025-04-141 source
GPT-4.1 nano benchmark scores: 10 benchmarks tracked. 40% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 10 benchmarks·best result 70 on MATH level 5 (reproduced)·4 of 10 independently reproduced·$0.1/$0.4 per M tokens
Source: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MATH level 5 | 70accuracy (%) | reproduced· optimizedT1 | 2025-04-14Epoch AI |
| GPQA diamond | 31.9accuracy (%) | reproduced· optimizedT1 | 2025-04-14Epoch AI |
| OTIS Mock AIME 2024-2025 | 28.82accuracy (%) | reproduced· optimizedT1 | 2025-04-14Epoch AI |
| Fiction.LiveBench | 25accuracy (%) | unverifiedT1 | 2025-04-14Fiction.live leaderboard |
| WeirdML | 18.98accuracy (%) | unverifiedT1 | 2025-04-14WeirdML Leaderboard |
| Aider polyglot | 8.9accuracy (%) | unverified· optimizedT1 | 2025-04-14Aider LLM Leaderboards |
| FrontierMath-2025-02-28-Private | 1.81accuracy (%) | reproduced· optimizedT1 | 2025-04-14Epoch AI |
| CritPt | 0accuracy (%) | unverifiedT2 | 2025-04-14 |
| ARC-AGI | 0accuracy (%) | unverified· optimizedT1 | 2025-04-14ARC Prize Leaderboard |
| ARC-AGI-2 | 0accuracy (%) | unverifiedT2 | 2025-04-14 |
Claims drawn from cited facts, not live model generation.
This model was developed by OpenAI and released in 2025.
2 cited facts
10 scored benchmarks is enough coverage to sketch a profile, not enough to settle arguments. None of its scores lead their tables, which limits what this record can support — presence in the measured field, not category leadership. Because the scores have been independently reproduced, outside evaluation backs them up, so the reader can trust the record.
3 cited facts
The model trails the state-of-the-art leader by an average of 66.06 points, placing it well behind the frontrunners. Others define the frontier here — zero leads says only that, leaving open the distance question that actually ranks a model.
2 cited facts
6 facts cross-checked across data sources: 3 corroborated, 2 single-source, 1 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, modelsdev_models, openrouter_models
GPT-4.1 nano is an AI model developed by OpenAI, released 2025-04-14. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-4.1 nano has recorded scores on 10 benchmarks, each shown with its evidence status.
GPT-4.1 nano has recorded scores on 10 benchmarks — GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, FrontierMath-2025-02-28-Private, Aider polyglot, and 4 more. The full table above shows each score with its evidence status.
4 of 10 recorded scores (40%) are independently reproduced rather than self-reported by the lab.
GPT-4.1 nano has 10 tracked claims: 4 independently reproduced, 6 unverified.
Listed API pricing: $0.1 per million input tokens, $0.4 per million output tokens. See the pricing block for the full breakdown.