Microsoft·released 2023-12-13✓ 2 sources
Phi-2 benchmark scores: 8 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 8 benchmarks·best result 67.87 on ARC AI2 (self-reported)·0 of 8 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| ARC AI2 | 67.87accuracy (%) | self-reported· optimizedT1 | 2023-12-12Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| OpenBookQA | 64.8accuracy (%) | self-reported· optimizedT1 | 2023-12-12Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| BBH | 45.87accuracy (%) | self-reported· optimizedT1 | 2023-12-12Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| TriviaQA | 45.2accuracy (%) | self-reported· optimizedT1 | 2023-12-12Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| MMLU | 44.53accuracy (%) | unverified· optimizedT1 | 2023-12-12Stanford CRFM Leaderboard |
| HellaSwag | 38.13accuracy (%) | self-reported· optimizedT1 | 2023-12-12Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| ANLI | 13.75accuracy (%) | self-reported· optimizedT1 | 2023-12-12Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| Winogrande | 9.4accuracy (%) | self-reported· optimizedT1 | 2023-12-12Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
Claims drawn from cited facts, not live model generation.
Developed by Microsoft, this model is a mid-size 2.78-billion-parameter system released in 2023; that scale keeps inference efficient and low cost while offering less headroom than larger models.
3 cited facts
This model is tracked across eight benchmarks but currently holds no top score in any of them. Because these results are self-reported, they should be treated as claims rather than independently verified.
3 cited facts
On the disclosed harness, it trails the average SOTA leader by 39.35 points, a wide gap that places it well behind the leaders. With zero SOTA leader positions across the measured benchmarks, it currently holds no category-leading score.
2 cited facts
2 facts cross-checked across data sources: 1 corroborated, 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks, huggingface_models
Phi-2 is an AI model developed by Microsoft, released 2023-12-13. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Phi-2 has recorded scores on 8 benchmarks, each shown with its evidence status.
Phi-2 has recorded scores on 8 benchmarks — OpenBookQA, Winogrande, TriviaQA, MMLU, ARC AI2, ANLI, and 2 more. The full table above shows each score with its evidence status.
0 of 8 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Phi-2 has 8 tracked claims: 7 self-reported, 1 unverified.