Shanghai AI Laboratory·released 2023-09-181 source
internlm-20b benchmark scores: 6 benchmarks tracked. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 0 of 6 benchmarks·best result 75.6 on ARC AI2 (self-reported)·0 of 6 independently reproduced
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| ARC AI2 | 75.6accuracy (%) | self-reported· optimizedT1 | 2023-09-18Qwen Technical Report |
| LAMBADA | 71.8accuracy (%) | self-reported· optimizedT1 | 2023-09-18Qwen Technical Report |
| HellaSwag | 70.8accuracy (%) | self-reported· optimizedT1 | 2023-09-18Qwen Technical Report |
| GSM8K | 62.9accuracy (%) | unverified· optimizedT1 | 2023-09-18Yi: Open Foundation Models by 01.AI |
| PIQA | 60.6accuracy (%) | self-reported· optimizedT1 | 2023-09-18Qwen Technical Report |
| BBH | 36.67accuracy (%) | self-reported· optimizedT1 | 2023-09-18Qwen Technical Report |
Claims drawn from cited facts, not live model generation.
This model was developed by Shanghai AI Laboratory and released in 2023.
2 cited facts
Interpret this thin record cautiously: 6 scored benchmarks without a lead neither ranks the model nor counts against it. These results are self-reported, meaning they are vendor-claimed and not independently confirmed, so they should be treated as claims rather than verified facts.
3 cited facts
With an average SOTA gap of 24.27 points, the model is well behind the leaders, indicating a wide performance deficit. Behind the leader on every tracked benchmark, by margins this count cannot see — position in the field is established, the size of the deficit is not.
2 cited facts
1 facts cross-checked across data sources: 1 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: epoch_benchmarks
internlm-20b is an AI model developed by Shanghai AI Laboratory, released 2023-09-18. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
internlm-20b has recorded scores on 6 benchmarks, each shown with its evidence status.
internlm-20b has recorded scores on 6 benchmarks — PIQA, ARC AI2, BBH, GSM8K, HellaSwag, LAMBADA. The full table above shows each score with its evidence status.
0 of 6 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
internlm-20b has 6 tracked claims: 5 self-reported, 1 unverified.