Microsoft·released 2024-04-231 source
phi-3-medium 14B benchmark scores: 10 benchmarks tracked. 20% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 10 benchmarks·best result 88.8 on ARC AI2 (self-reported)·2 of 10 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| ARC AI2 | 88.8accuracy (%) | self-reported· optimizedT1 | 2024-04-23Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| OpenBookQA | 83.2accuracy (%) | self-reported· optimizedT1 | 2024-04-23Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| HellaSwag | 76.53accuracy (%) | self-reported· optimizedT1 | 2024-04-23Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| BBH | 75.2accuracy (%) | self-reported· optimizedT1 | 2024-04-23Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| TriviaQA | 73.9accuracy (%) | self-reported· optimizedT1 | 2024-04-23Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| MMLU | 70.67accuracy (%) | unverified· optimizedT1 | 2024-04-23Stanford CRFM Leaderboard |
| Winogrande | 63accuracy (%) | self-reported· optimizedT1 | 2024-04-23Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| ANLI | 33.7accuracy (%) | self-reported· optimizedT1 | 2024-04-23Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| MATH level 5 | 17.56accuracy (%) | reproduced· optimizedT1 | 2024-04-23Epoch AI |
| GPQA diamond | 3.45accuracy (%) | reproduced· optimizedT1 | 2024-04-23Epoch AI |
Claims drawn from cited facts, not live model generation.
This model was developed by Microsoft and released in 2024.
2 cited facts
Across 10 tracked benchmarks this model places without leading — coverage wide enough that the off-the-front reading is earned rather than accidental. Because the record is independently reproduced, outside evaluation backs these scores up.
3 cited facts
This model trails the SOTA leader on average by 24.93 points, a wide gap that places it well behind the frontier. It holds the top score on zero benchmarks, meaning it has no genuine category leadership.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
phi-3-medium 14B is an AI model developed by Microsoft, released 2024-04-23. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
phi-3-medium 14B has recorded scores on 10 benchmarks, each shown with its evidence status.
phi-3-medium 14B has recorded scores on 10 benchmarks — OpenBookQA, Winogrande, TriviaQA, GPQA diamond, MATH level 5, MMLU, and 4 more. The full table above shows each score with its evidence status.
2 of 10 recorded scores (20%) are independently reproduced rather than self-reported by the lab.
phi-3-medium 14B has 10 tracked claims: 2 independently reproduced, 7 self-reported, 1 unverified.