IBM Research·released 2024-08-141 source
PowerMoE-3b benchmark scores: 11 benchmarks tracked, leading 9. 0% of its scores are independently reproduced — the rest are self-reported or unverified.
Leads on 9 of 11 benchmarks·best result 79.1 on PIQA (unverified)·0 of 11 independently reproduced
Leads on: ARC, BoolQ, GSM8k (5 shot), Hellaswag, MBPP, MMLU (5 shot), PIQA, humaneval, math (4 shot)
| Benchmark | Score | Evidence status | Measured |
|---|---|---|---|
| PIQA | 79.1accuracy-norm | unverifiedT2 | — |
| Hellaswag | 71.5accuracy-norm | unverifiedT2 | — |
| BoolQ | 65accuracy | unverifiedT2 | — |
| Winogrande | 65accuracy-norm | unverifiedT2 | — |
| ARC | 58.1accuracy-norm | unverifiedT2 | — |
| MMLU (5 shot) | 42.8accuracy | unverifiedT2 | — |
| OpenBookQA | 41accuracy-norm | unverifiedT2 | — |
| MBPP | 32.4pass@1 | unverifiedT2 | — |
| GSM8k (5 shot) | 25.9accuracy | unverifiedT2 | — |
| humaneval | 20.1pass@1 | unverifiedT2 | — |
| math (4 shot) | 14.8accuracy | unverifiedT2 | — |
Claims drawn from cited facts, not live model generation.
Developed by IBM Research and released in 2024, this model is a compact-scale system with 3.37 billion parameters, trading raw headroom for greater efficiency and lower serving cost.
3 cited facts
This model is evaluated across 11 tracked benchmarks, and it currently holds the top score on 9 of them. Because the record is unverified, these numbers are neither independently confirmed nor disputed, so they should be treated as unconfirmed claims rather than established results.
3 cited facts
On average the model trails the top benchmark score by 5.13 points under the disclosed harness, a moderate gap that reads as competitive but not front-of-pack. It holds the top score on 9 benchmarks, yet this breadth of leadership is tempered by the overall average shortfall.
3 cited facts
2 facts cross-checked across data sources: 2 single-source.
Source facts, citations, and refresh stamp for this record.
Sources: huggingface_models
PowerMoE-3b is an AI model developed by IBM Research, released 2024-08-14. Its benchmark record below tags every score as independently reproduced, self-reported, or unverified.
PowerMoE-3b has recorded scores on 11 benchmarks, each shown with its evidence status.
PowerMoE-3b has recorded scores on 11 benchmarks — ARC, BoolQ, Hellaswag, OpenBookQA, PIQA, Winogrande, and 5 more. The full table above shows each score with its evidence status.
0 of 11 recorded scores (0%) are independently reproduced rather than self-reported by the lab.
PowerMoE-3b has 11 tracked claims: 11 unverified.