Prime Intellect·released 2024-11-291 source
INTELLECT-1 benchmark scores: 6 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 6 benchmarks·best result 61.89 on HellaSwag (self-reported)·0 of 6 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| HellaSwag | 61.89accuracy (%) | self-reported· optimizedT1 | 2024-11-29INTELLECT-1 Technical Report |
| ARC AI2 | 39.36accuracy (%) | self-reported· optimizedT1 | 2024-11-29INTELLECT-1 Technical Report |
| GSM8K | 38.58accuracy (%) | self-reported· optimizedT1 | 2024-11-29INTELLECT-1 Technical Report |
| MMLU | 33.19accuracy (%) | self-reported· optimizedT1 | 2024-11-29INTELLECT-1 Technical Report |
| Winogrande | 31.64accuracy (%) | self-reported· optimizedT1 | 2024-11-29INTELLECT-1 Technical Report |
| BBH | 13.13accuracy (%) | self-reported· optimizedT1 | 2024-11-29INTELLECT-1 Technical Report |
Claims drawn from cited facts, not live model generation.
This model was developed by Prime Intellect in 2024.
2 cited facts
Small record, no leads — at 6 tracked scores the counts are context; whatever signal exists sits in the individual cells. Because these results are self-reported, they should be treated as claims rather than verified results.
3 cited facts
The model trails the state-of-the-art by an average of 51.63 points across benchmarks, and it holds the top score on none of them, indicating a wide gap behind the leaders.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
INTELLECT-1 is an AI model developed by Prime Intellect, released 2024-11-29. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
INTELLECT-1 has recorded scores on 6 benchmarks, each shown with its evidence status.
INTELLECT-1 has recorded scores on 6 benchmarks — Winogrande, MMLU, ARC AI2, BBH, GSM8K, HellaSwag. The full table above shows each score with its evidence status.
0 of 6 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
INTELLECT-1 has 6 tracked claims: 6 self-reported.