Visual Physics Comprehension Test
Integrity rank #35 of 61 · 26 models scored · top score 86.5 · Gemini 3 Pro
Visual/spatial physical intuition - trajectory prediction from a static diagram
Strongest on discrimination, weakest on contamination resistance. Caveat: the public test set does not rule out training-data contamination.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | Gemini 3 Pro | 86.5 | unverifiedT1 | 2025-11-18 |
| 2 | GPT-5.2 | 76 | unverifiedT1 | 2025-12-11 |
| 3 | Gemini 3 Flash | 58.9 | unverifiedT1 | 2025-12-17 |
| 4 | GPT-5 | 49 | unverifiedT1 | 2025-08-07 |
| 5 | GPT-5.1 | 38.05 | unverifiedT1 | 2025-11-13 |
| 6 | o4-mini | 36.25 | unverifiedT1 | 2025-04-16 |
| 7 | o3 | 28 | unverifiedT1 | 2024-12-20 |
| 8 | Gemini 2.5 Pro (Mar 2025) | 22 | unverifiedT1 | 2025-03-25 |
| 9 | Gemini 2.5 Pro (Jun 2025) | 19.6 | unverifiedT1 | 2025-06-05 |
| 10 | Gemini 2.5 Flash (Sep 2025) | 19.3 | unverifiedT1 | 2025-09-25 |
| 11 | GPT-4.5 | 17.5 | unverifiedT1 | 2025-02-27 |
| 12 | Gemini 2.5 Pro (May 2025) | 10.75 | unverifiedT1 | 2025-05-06 |
| 13 | GPT-5 mini | 10.3 | unverifiedT1 | 2025-08-07 |
| 14 | Claude Opus 4.5 | 10 | unverifiedT1 | 2025-11-24 |
| 15 | GPT-4o (Nov 2024) | 10 | unverifiedT1 | 2024-05-13 |
| 16 | Claude Sonnet 4.5 | 9.7 | unverifiedT1 | 2025-09-29 |
| 17 | Claude 3.7 Sonnet | 8.5 | unverifiedT1 | 2025-02-24 |
| 18 | Gemini 2.5 Flash (Apr 2025) | 7 | unverifiedT1 | 2025-04-17 |
| 19 | Claude Opus 4 | 7 | unverifiedT1 | 2025-05-22 |
| 20 | GPT-5 nano | 5.8 | unverifiedT1 | 2025-08-07 |
| 21 | o1 | 5.5 | unverifiedT1 | 2024-12-05 |
| 22 | Claude Opus 4.1 | 2.5 | unverifiedT1 | 2025-08-05 |
| 23 | Claude Sonnet 4 | 1 | unverifiedT1 | 2025-05-22 |
| 24 | GPT-4o mini | 1 | unverifiedT1 | 2024-07-18 |
| 25 | Claude 3.5 Sonnet | 0 | unverifiedT1 | 2024-06-20 |
| 26 | Claude 3.5 Sonnet (October 2024) | 0 | unverifiedT1 | 2024-10-22 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
This benchmark scores 26 models, with a score spread of 86.5 points, indicating a wide spread that supports genuine capability separation between models.
2 cited facts
The benchmark's Integrity Index of 87 (grade A) and rank of 26 out of 51 reflect its health as a discriminator of frontier models under a disclosed harness. Its weakest component is contamination, which narrows how much to trust top scores.
4 cited facts
No crowd at the top here — a single model holds 86.5 with 13.5 points of ceiling to spare, so the lead reads as clean separation and the benchmark still has room to record a successor. Because the benchmark is not saturated, there is still genuine room to separate the field at the top, so small differences in ranking reflect real capability gaps.
5 cited facts
Scores from this benchmark are directly comparable as the evaluation harness is consistent, but because the test set is public, high scores warrant increased scrutiny for potential contamination. Contamination risk is considered low due to procedurally generated images, but no formal contamination analysis has been reported.
3 cited facts
This benchmark measures visual-spatial physical intuition by requiring trajectory prediction from static diagrams. Procedural generation with simulator-checked answers gives an unlimited, objectively graded item supply. The trade is narrowness: every task is a variation on the same falling-ball scene. If the real failure lies somewhere other than visual comprehension, the score misattributes it; the single-answer format also assumes away physically ambiguous scenes. Both assumptions ride along silently with every number. The sharpest caveat is that the benchmark's small size, narrow single-task focus, and very small human sample limit its reliability and generalizability.
4 cited facts
VPCT (Visual Physics Comprehension Test): Visual Physics Comprehension Test: a multimodal benchmark of 100 diagram problems, each showing a floating ball, a series of ramps, and three buckets; the model must predict which bucket the ball falls into. It probes basic physical intuition that humans find trivial but models struggle with. Included in Epoch AI's Capabilities Index but authored by a third party.
Gemini 3 Pro leads VPCT at 86.5. The full leaderboard above lists every recorded measurement, not just the headline number.
26 models have recorded scores on VPCT, spanning a score spread of 86.5.
tensor.news grades VPCT A for integrity (score 86/100), ranking #35 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — VPCT still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is public, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Scores are reported under a consistent harness, so comparisons on VPCT are reasonably apples-to-apples.
Every VPCT measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites epoch.ai/benchmarks/vpct.