Microsoft·released 2024-12-11✓ 2 sources
Phi-4 benchmark scores: 6 benchmarks tracked. 50% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 6 benchmarks·best result 79.73 on MMLU (self-reported)·3 of 6 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| MMLU | 79.73accuracy (%) | self-reported· optimizedT1 | 2024-12-12Phi-4 Technical Report |
| MATH level 5 | 64.94accuracy (%) | reproduced· optimizedT1 | 2024-12-12Epoch AI |
| Lech Mazur Writing | 62.6accuracy (%) | unverifiedT1 | 2024-12-12lechmazur/writing Github repository |
| GPQA diamond | 41.41accuracy (%) | reproduced· optimizedT1 | 2024-12-12Epoch AI |
| OTIS Mock AIME 2024-2025 | 13.66accuracy (%) | reproduced· optimizedT1 | 2024-12-12Epoch AI |
| Balrog | 11.6accuracy (%) | unverifiedT1 | 2024-12-12Balrog Leaderboard |
Claims drawn from cited facts, not live model generation.
Created by Microsoft and released in 2024, this model is a mid-size system with 14.66 billion parameters. That scale balances efficiency and cost-effectiveness with enough headroom for many tasks, though it lacks the raw capacity of frontier-scale models.
4 cited facts
This model is tracked on 6 benchmarks and currently holds no top score on any of them. Because its results have been independently reproduced by outside evaluation, this record can be treated as verified rather than as mere vendor claims.
3 cited facts
Relative to disclosed leaderboard scores, this model trails the top score by an average of 40.92 points under the given harness, a wide gap that places it well behind the front of the pack. It holds the top score on none of the tracked state-of-the-art benchmarks (0 of them), so it does not yet show genuine category leadership.
2 cited facts
6 facts cross-checked across data sources: 1 corroborated, 5 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, huggingface_models, openrouter_models
Phi-4 is an AI model developed by Microsoft, released 2024-12-11. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Phi-4 has recorded scores on 6 benchmarks, each shown with its evidence status.
Phi-4 has recorded scores on 6 benchmarks — Lech Mazur Writing, GPQA diamond, MATH level 5, OTIS Mock AIME 2024-2025, MMLU, Balrog. The full table above shows each score with its evidence status.
3 of 6 recorded scores (50%) are independently reproduced rather than self-reported by the lab.
Phi-4 has 6 tracked claims: 3 independently reproduced, 1 self-reported, 2 unverified.