Salesforce AI Research·released 2023-09-071 source
XGen-7B benchmark scores: 6 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 6 benchmarks·best result 65.6 on HellaSwag (self-reported)·0 of 6 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| HellaSwag | 65.6accuracy (%) | self-reported· optimizedT1 | 2023-09-07XGen-7B Technical Report |
| PIQA | 51accuracy (%) | self-reported· optimizedT1 | 2023-09-07XGen-7B Technical Report |
| Winogrande | 29.8accuracy (%) | self-reported· optimizedT1 | 2023-09-07XGen-7B Technical Report |
| ARC AI2 | 21.6accuracy (%) | self-reported· optimizedT1 | 2023-09-07XGen-7B Technical Report |
| OpenBookQA | 20.27accuracy (%) | self-reported· optimizedT1 | 2023-09-07XGen-7B Technical Report |
| MMLU | 15.07accuracy (%) | self-reported· optimizedT1 | 2023-09-07XGen-7B Technical Report |
Claims drawn from cited facts, not live model generation.
This model was developed by Salesforce AI Research and released in 2023.
2 cited facts
The base here is 6 tracked benchmarks — modest, so give the spread of results more credence than any headline number. Among its measured boards the top slot goes elsewhere every time; a rank observation, and only that. Because the results are self-reported, the numbers are vendor-claimed and not yet independently confirmed, so read them as claims rather than verified results.
3 cited facts
This model trails the state-of-the-art leader by an average of 51.62 points across evaluated benchmarks, a wide gap that places it well behind the leading models.
1 cited fact
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
XGen-7B is an AI model developed by Salesforce AI Research, released 2023-09-07. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
XGen-7B has recorded scores on 6 benchmarks, each shown with its evidence status.
XGen-7B has recorded scores on 6 benchmarks — PIQA, OpenBookQA, Winogrande, MMLU, ARC AI2, HellaSwag. The full table above shows each score with its evidence status.
0 of 6 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
XGen-7B has 6 tracked claims: 6 self-reported.