LMSYS·released 2023-04-121 source
vicuna-13b-v1.1 benchmark scores: 7 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 7 benchmarks·best result 54.8 on PIQA (self-reported)·0 of 7 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| PIQA | 54.8accuracy (%) | self-reported· optimizedT1 | 2023-04-12Textbooks Are All You Need II: phi-1.5 technical report |
| HellaSwag | 43.73accuracy (%) | self-reported· optimizedT1 | 2023-04-12Textbooks Are All You Need II: phi-1.5 technical report |
| Winogrande | 41.6accuracy (%) | self-reported· optimizedT1 | 2023-04-12Textbooks Are All You Need II: phi-1.5 technical report |
| GSM8K | 28.13accuracy (%) | unverified· optimizedT1 | 2023-04-12Baichuan 2: Open Large-scale Language Models |
| ARC AI2 | 24.27accuracy (%) | self-reported· optimizedT1 | 2023-04-12Textbooks Are All You Need II: phi-1.5 technical report |
| BBH | 24.05accuracy (%) | unverified· optimizedT1 | 2023-04-12Baichuan 2: Open Large-scale Language Models |
| OpenBookQA | 10.67accuracy (%) | self-reported· optimizedT1 | 2023-04-12Textbooks Are All You Need II: phi-1.5 technical report |
Claims drawn from cited facts, not live model generation.
This model was developed by LMSYS and released in 2023.
2 cited facts
Present on 7 tracked benchmarks and topping none, this model reads as part of the field — a standing worth knowing that stops short of a verdict. Since its scores are self-reported, they should be read as claims rather than verified results.
3 cited facts
With a 54.19-point average shortfall, the leaders here are effectively out of frame — a wide gap under the disclosed harnesses, and a blended figure whose weight depends on which benchmarks feed it.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
vicuna-13b-v1.1 is an AI model developed by LMSYS, released 2023-04-12. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
vicuna-13b-v1.1 has recorded scores on 7 benchmarks, each shown with its evidence status.
vicuna-13b-v1.1 has recorded scores on 7 benchmarks — PIQA, OpenBookQA, Winogrande, ARC AI2, BBH, GSM8K, and 1 more. The full table above shows each score with its evidence status.
0 of 7 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
vicuna-13b-v1.1 has 7 tracked claims: 5 self-reported, 2 unverified.