Physical Interaction: Question Answering
Integrity rank #57 of 61 · 44 models scored · top score 79.1 · PowerMoE-3b
Physical commonsense reasoning about everyday goals and solutions
Strongest on discrimination, weakest on contamination resistance. Caveat: scores come from an inconsistent mix of harnesses, so head-to-head comparisons are not apples-to-apples.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | PowerMoE-3b | 79.1 | unverifiedT2 | — |
| 2 | GPT-4o mini | 77.4 | self-reported· optimizedT1 | 2024-07-18 |
| 3 | Gemini 1.5 Flash (Sep 2024) | 75 | self-reported· optimizedT1 | 2024-05-10 |
| 4 | Llama 3.1-405B | 71.8 | self-reported· optimizedT1 | 2024-07-23 |
| 5 | Falcon-180B | 69.8 | unverified· optimizedT1 | 2023-09-06 |
| 6 | DeepSeek-V3 | 69.4 | self-reported· optimizedT1 | 2024-12-24 |
| 7 | DeepSeek-V2 (MoE-236B, May 2024) | 67.8 | self-reported· optimizedT1 | 2024-05-07 |
| 8 | Gemma 2 9B | 67.4 | self-reported· optimizedT1 | 2024-06-24 |
| 9 | Mixtral 8x7B | 67.2 | unverified· optimizedT1 | 2023-12-11 |
| 10 | Mistral NeMo | 67 | self-reported· optimizedT1 | 2024-07-18 |
| 11 | Stable Beluga 2 | 66.6 | self-reported· optimizedT1 | 2023-07-20 |
| 12 | Mistral 7B v0.1 | 66 | self-reported· optimizedT1 | 2023-10-10 |
| 13 | Falcon-40B | 66 | unverified· optimizedT1 | 2023-03-15 |
| 14 | LLaMA-65B | 65.6 | self-reported· optimizedT1 | 2023-02-24 |
| 15 | Llama 2-70B | 65.6 | self-reported· optimizedT1 | 2023-07-18 |
| 16 | Qwen2.5-72B | 65.2 | self-reported· optimizedT1 | 2024-09-19 |
| 17 | Nemotron-4 15B | 64.8 | self-reported· optimizedT1 | 2024-02-27 |
| 18 | LLaMA-33B | 64.6 | self-reported· optimizedT1 | 2023-02-27 |
| 19 | Llama 2-34B | 63.8 | unverified· optimizedT1 | 2023-07-18 |
| 20 | MPT-30B | 63.8 | unverified· optimizedT1 | 2023-06-22 |
| 21 | Llama 3.1-8B | 62.4 | self-reported· optimizedT1 | 2024-07-23 |
| 22 | Gemma 7B | 62.4 | unverified· optimizedT1 | 2024-02-21 |
| 23 | Llama 2-13B | 61.6 | unverified· optimizedT1 | 2023-07-18 |
| 24 | MPT-7B | 61.2 | self-reported· optimizedT1 | 2023-05-05 |
| 25 | Falcon-7B | 60.6 | self-reported· optimizedT1 | 2023-04-24 |
| 26 | internlm-20b | 60.6 | self-reported· optimizedT1 | 2023-09-18 |
| 27 | LLaMA-13B | 60.2 | self-reported· optimizedT1 | 2023-02-27 |
| 28 | Qwen-14B | 59.8 | self-reported· optimizedT1 | 2023-09-24 |
| 29 | LLaMA-7B | 59.6 | self-reported· optimizedT1 | 2023-02-24 |
| 30 | Llama 2-7B | 57.6 | unverified· optimizedT1 | 2023-07-18 |
| 31 | Baichuan2-13B | 56.2 | self-reported· optimizedT1 | 2023-09-06 |
| 32 | Qwen-7B | 55.8 | self-reported· optimizedT1 | 2023-09-28 |
| 33 | internlm-7b | 55.8 | self-reported· optimizedT1 | 2023-07-05 |
| 34 | vicuna-13b-v1.1 | 54.8 | self-reported· optimizedT1 | 2023-04-12 |
| 35 | Gemma 2B | 54.6 | unverified· optimizedT1 | 2024-02-21 |
| 36 | RedPajama-INCITE-7B-Base | 53.8 | self-reported· optimizedT1 | 2023-05-04 |
| 37 | Baichuan1-7B | 52.4 | self-reported· optimizedT1 | 2023-06-01 |
| 38 | open_llama_7b | 52 | self-reported· optimizedT1 | 2023-06-07 |
| 39 | XGen-7B | 51 | self-reported· optimizedT1 | 2023-09-07 |
| 40 | Dolly 2.0-12b | 50.8 | self-reported· optimizedT1 | 2023-04-12 |
| 41 | Cerebras-GPT-13B | 47 | self-reported· optimizedT1 | 2023-04-06 |
| 42 | Qwen-1_8B | 46.6 | self-reported· optimizedT1 | 2023-11-30 |
| 43 | chatglm2-6b | 39.2 | self-reported· optimizedT1 | 2023-06-24 |
| 44 | stablelm-tuned-alpha-7b | 31.6 | self-reported· optimizedT1 | 2023-04-19 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
Mid-band on both counts, 44 models with 47.5 points of spread; big gaps here are believable, adjacent ranks less so. Such a wide spread indicates that the differences between model ranks are meaningful and reflect genuine capability gaps.
3 cited facts
The Benchmark Integrity Index is 62 (grade C), ranking 47th out of 51 benchmarks. The weakest component is contamination, which narrows how much to trust top scores; this index reflects the benchmark's health as a discriminator of frontier models under a disclosed harness.
5 cited facts
The benchmark is not saturated, with a top score of 79.1 and a headroom of 20.9 points ; only two models cluster near the top within a couple of points, meaning there is genuine room to distinguish models at the top rather than noise.
4 cited facts
Under its disclosed harness, which exhibits label order and normalization sensitivity, this benchmark's scores are not directly comparable across different implementations, meaning they should not be compared head-to-head. Because the test set is public and has been available since 2019, high scores deserve additional scrutiny for possible contamination, especially given that the benchmark has been treated as contaminated in common pretraining corpora.
4 cited facts
A probe is not a guarantee: scores index performance on everyday physical scenarios as posed in the test, and transfer to acting in the physical world is exactly what the format cannot show. It was built from binary QA pairs crowdsourced from instructables.com, with adversarial filtering to reduce trivial patterns. Choice among given solutions is weaker evidence than producing a solution; a model can prefer the workable option without being able to generate it, so the score bounds the skill from above. A key caveat: crowd-authored scenarios can be underspecified or ambiguous, and known label noise and instructables-domain skew limit how far performance generalizes to other physical contexts.
4 cited facts
PIQA (Physical Interaction: Question Answering): Physical Interaction: Question Answering — a binary-choice physical-commonsense benchmark: given a goal, choose the workable everyday-physical solution over a near-miss distractor; scored as accuracy.
PowerMoE-3b leads PIQA at 79.1. The full leaderboard above lists every recorded measurement, not just the headline number.
44 models have recorded scores on PIQA, spanning a score spread of 47.5.
tensor.news grades PIQA C for integrity (score 62/100), ranking #57 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — PIQA still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is public, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Only partly — scores come from a mix of harnesses and protocols, so a head-to-head on PIQA is not fully apples-to-apples.
Every PIQA measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites arxiv.org/abs/1911.11641.