Microsoft·released 2024-04-231 source
phi-3-small 7.4B benchmark scores: 8 benchmarks tracked, leading 2. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 2 of 8 benchmarks·best result 87.6 on ARC AI2 (self-reported)·0 of 8 independently reproduced
Head-to-headphi-3-small 7.4B vs GPT-3.5 Turbo (Nov 2023)Leads on: ANLI, OpenBookQA
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| ARC AI2 | 87.6accuracy (%) | self-reported· optimizedT1 | 2024-04-23Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| OpenBookQA | 84accuracy (%) | self-reported· optimizedT1 | 2024-04-23Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| BBH | 72.13accuracy (%) | self-reported· optimizedT1 | 2024-04-23Phi-3 Technical Report: A Highly Capable Language Model Locally on Your PhonePhi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| HellaSwag | 69.33accuracy (%) | self-reported· optimizedT1 | 2024-04-23Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| MMLU | 67.6accuracy (%) | self-reported· optimizedT1 | 2024-04-23Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| Winogrande | 63accuracy (%) | self-reported· optimizedT1 | 2024-04-23Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| TriviaQA | 58.1accuracy (%) | self-reported· optimizedT1 | 2024-04-23Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
| ANLI | 37.15accuracy (%) | self-reported· optimizedT1 | 2024-04-23Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone |
Claims drawn from cited facts, not live model generation.
This model was developed by Microsoft. It was released in 2024.
2 cited facts
Holding top scores on 2 of 8 tracked benchmarks gives this model genuine but narrow leadership — worth checking whether those leads cover the tasks you care about. Because the record is self-reported, the numbers are vendor-claimed and not yet independently confirmed; they should be read as claims, not verified results.
3 cited facts
With an average gap of 13.18 points to the leader, this model is well behind the state of the art, though it does top 2 benchmarks.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: curated_capability_claims, epoch_benchmarks
phi-3-small 7.4B is an AI model developed by Microsoft, released 2024-04-23. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
phi-3-small 7.4B has recorded scores on 8 benchmarks, each shown with its evidence status.
phi-3-small 7.4B has recorded scores on 8 benchmarks — OpenBookQA, Winogrande, TriviaQA, MMLU, ARC AI2, ANLI, and 2 more. The full table above shows each score with its evidence status.
0 of 8 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
phi-3-small 7.4B has 8 tracked claims: 8 self-reported.