Microsoft·released 2023-09-111 source
Phi-1.5 benchmark scores: 5 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 5 benchmarks·best result 46.8 on Winogrande (self-reported)·0 of 5 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| Winogrande | 46.8accuracy (%) | self-reported· optimizedT1 | 2023-09-11Textbooks Are All You Need II: phi-1.5 technical report |
| HellaSwag | 30.13accuracy (%) | self-reported· optimizedT1 | 2023-09-11Textbooks Are All You Need II: phi-1.5 technical report |
| ARC AI2 | 25.87accuracy (%) | self-reported· optimizedT1 | 2023-09-11Textbooks Are All You Need II: phi-1.5 technical report |
| MMLU | 16.8accuracy (%) | self-reported· optimizedT1 | 2023-09-11Textbooks Are All You Need II: phi-1.5 technical report |
| OpenBookQA | 16.27accuracy (%) | self-reported· optimizedT1 | 2023-09-11Textbooks Are All You Need II: phi-1.5 technical report |
Claims drawn from cited facts, not live model generation.
This model was developed by Microsoft and released in 2023.
2 cited facts
Its record spans 5 tracked benchmarks — a narrow slice, so weight the pattern across them more than any single number. None of those results is a top finish, which locates the model in the field and supports nothing stronger. The record is self-reported, meaning the numbers are vendor-claimed and not independently confirmed, so they should be treated as claims rather than verified results.
3 cited facts
This model trails the state-of-the-art leader by an average of 59.62 points, a wide gap that places it well behind the leaders, and it holds the top score on none of the benchmarks.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Phi-1.5 is an AI model developed by Microsoft, released 2023-09-11. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Phi-1.5 has recorded scores on 5 benchmarks, each shown with its evidence status.
Phi-1.5 has recorded scores on 5 benchmarks — OpenBookQA, Winogrande, MMLU, ARC AI2, HellaSwag. The full table above shows each score with its evidence status.
0 of 5 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Phi-1.5 has 5 tracked claims: 5 self-reported.