OpenAI·released 2024-07-181 source
GPT-4o mini benchmark scores: 15 benchmarks tracked. 27% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 15 benchmarks·best result 91.3 on GSM8K (self-reported)·4 of 15 independently reproduced·$0.15/$0.6 per M tokens
Source: models.dev, OpenRouter · See every model’s pricing →
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| GSM8K | 91.3accuracy (%) | self-reported· optimizedT1 | 2024-07-18Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| PIQA | 77.4accuracy (%) | self-reported· optimizedT1 | 2024-07-18Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| MMLU | 75.73accuracy (%) | self-reported· optimizedT1 | 2024-07-18Phi-4 Technical Report |
| Lech Mazur Writing | 67.2accuracy (%) | unverifiedT1 | 2024-07-18lechmazur/writing Github repository |
| GeoBench | 64accuracy (%) | unverifiedT1 | 2024-07-18GeoBench leaderboard |
| MATH level 5 | 52.63accuracy (%) | reproduced· optimizedT1 | 2024-07-18Epoch AI |
| Balrog | 17.4accuracy (%) | unverifiedT1 | 2024-07-18Balrog Leaderboard |
| GPQA diamond | 16.96accuracy (%) | reproduced· optimizedT1 | 2024-07-18Epoch AI |
| WeirdML | 11.76accuracy (%) | unverifiedT1 | 2024-07-18WeirdML Leaderboard |
| OTIS Mock AIME 2024-2025 | 6.85accuracy (%) | reproduced· optimizedT1 | 2024-07-18Epoch AI |
| Aider polyglot | 3.6accuracy (%) | unverified· optimizedT1 | 2024-07-18Aider LLM Leaderboards |
| VPCT | 1accuracy (%) | unverifiedT1 | 2024-07-18VPCT leaderboard |
| Chess Puzzles | 0accuracy (%) | reproducedT1 | 2024-07-18Epoch AI |
| SimpleBench | 0accuracy (%) | unverifiedT1 | 2024-07-18SimpleBench Leaderboard |
| ARC-AGI-2 | 0accuracy (%) | unverifiedT2 | 2024-07-18 |
Claims drawn from cited facts, not live model generation.
Developed by OpenAI, this model was released in 2024.
2 cited facts
This model is evaluated on 15 tracked benchmarks, and it currently holds no top score on any of them. The results have been independently reproduced, so outside evaluation supports the record and the scores can be treated as verified.
3 cited facts
With an average SOTA gap of 52.8 points, it trails the leader by a wide margin and sits well behind the front-runners. It holds the top score on none of the SOTA leaderboard benchmarks, so it has no category-leading result yet.
2 cited facts
6 facts cross-checked across data sources: 3 corroborated, 2 single-source, 1 disputed.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, modelsdev_models, openrouter_models
GPT-4o mini is an AI model developed by OpenAI, released 2024-07-18. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
GPT-4o mini has recorded scores on 15 benchmarks, each shown with its evidence status.
GPT-4o mini has recorded scores on 15 benchmarks — Lech Mazur Writing, PIQA, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, WeirdML, and 9 more. The full table above shows each score with its evidence status.
4 of 15 recorded scores (27%) are independently reproduced rather than self-reported by the lab.
GPT-4o mini has 15 tracked claims: 4 independently reproduced, 3 self-reported, 8 unverified.
Listed API pricing: $0.15 per million input tokens, $0.6 per million output tokens. See the pricing block for the full breakdown.