Stability AI·released 2023-07-201 source
Stable Beluga 2 benchmark scores: 7 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 7 benchmarks·best result 81.47 on ARC AI2 (self-reported)·0 of 7 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| ARC AI2 | 81.47accuracy (%) | self-reported· optimizedT1 | 2023-07-20Qwen Technical Report |
| HellaSwag | 78.8accuracy (%) | self-reported· optimizedT1 | 2023-07-20Qwen Technical Report |
| LAMBADA | 71.3accuracy (%) | self-reported· optimizedT1 | 2023-07-20Qwen Technical Report |
| GSM8K | 69.6accuracy (%) | self-reported· optimizedT1 | 2023-07-20Qwen Technical Report |
| PIQA | 66.6accuracy (%) | self-reported· optimizedT1 | 2023-07-20Qwen Technical Report |
| BBH | 59.07accuracy (%) | self-reported· optimizedT1 | 2023-07-20Qwen Technical Report |
| MMLU | 58.13accuracy (%) | self-reported· optimizedT1 | 2023-07-20Qwen Technical Report |
Claims drawn from cited facts, not live model generation.
This model was developed by Stability AI and released in 2023.
2 cited facts
Its measured footprint is 7 benchmarks — partial coverage, so treat the aggregate picture as provisional. Never at the top across those boards, the model sits with the field; that calibrates expectations without capping them. Because the scores are self-reported, readers should consider them as claims rather than independently verified results.
3 cited facts
The model trails the state-of-the-art leader by an average of 17.59 points across benchmarks, a wide gap that places it well behind the front of the pack. It holds the top score on none of the evaluated benchmarks, meaning it has no genuine category leadership.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
Stable Beluga 2 is an AI model developed by Stability AI, released 2023-07-20. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
Stable Beluga 2 has recorded scores on 7 benchmarks, each shown with its evidence status.
Stable Beluga 2 has recorded scores on 7 benchmarks — PIQA, MMLU, ARC AI2, BBH, GSM8K, HellaSwag, and 1 more. The full table above shows each score with its evidence status.
0 of 7 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
Stable Beluga 2 has 7 tracked claims: 7 self-reported.