Complex Research using Integrated Thinking - Physics Test (CritPt)
Integrity rank #23 of 61 · 76 models scored · top score 32.3 · GPT-5.6 Sol
Frontier, research-level physics reasoning on novel unpublished problems
Strongest on contamination resistance, weakest on discrimination.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | GPT-5.6 Sol | 32.3 | unverifiedT2 | 2026-07-09 |
| 2 | GPT-5.5 Pro | 30.57 | unverifiedT2 | 2026-04-23 |
| 3 | GPT-5.4 Pro | 30 | unverifiedT2 | 2026-03-05 |
| 4 | GPT-5.6 Terra | 30 | unverifiedT2 | 2026-07-09 |
| 5 | Claude Opus 5 | 29.14 | unverifiedT2 | 2026-07-24 |
| 6 | Claude Fable 5 | 28.57 | unverifiedT2 | 2026-06-09 |
| 7 | GPT-5.5 | 27.14 | unverifiedT2 | 2026-04-23 |
| 8 | GPT-5.4 | 23.43 | unverifiedT2 | 2026-03-05 |
| 9 | Kimi K3 | 23.4 | unverifiedT2 | 2026-07-16 |
| 10 | GLM-5.2 | 20.86 | unverifiedT2 | 2026-06-16 |
| 11 | Claude Opus 4.8 | 20.86 | unverifiedT2 | 2026-05-28 |
| 12 | GPT-5.6 Luna | 20.6 | unverifiedT2 | 2026-07-09 |
| 13 | Gemini 3.1 Pro | 17.71 | unverifiedT2 | 2026-02-19 |
| 14 | Claude Sonnet 5 | 16.86 | unverifiedT2 | 2026-06-30 |
| 15 | DeepSeek V4 Flash 0731 | 16.57 | unverifiedT2 | 2026-07-31 |
| 16 | Grok 4.5 | 15.43 | unverifiedT2 | 2026-07-08 |
| 17 | Muse Spark 1.1 | 15.1 | unverifiedT2 | 2026-07-09 |
| 18 | Qwen3.7-Max | 13.43 | unverifiedT2 | 2026-05-19 |
| 19 | Gemini 3.5 Flash | 13.14 | unverifiedT2 | 2026-05-19 |
| 20 | DeepSeek-V4-Pro | 12.86 | unverifiedT2 | 2026-04-24 |
| 21 | GPT-5 | 12.6 | unverifiedT2 | 2025-08-07 |
| 22 | Claude Opus 4.7 | 12 | unverifiedT2 | 2026-04-16 |
| 23 | Muse Spark | 11.33 | unverifiedT2 | 2026-04-08 |
| 24 | Gemini 3.6 Flash | 10.57 | unverifiedT2 | 2026-07-21 |
| 25 | Kimi K2.7 Code | 10 | unverifiedT2 | 2026-06-12 |
| 26 | GPT-5.4 Mini | 10 | unverifiedT2 | 2026-03-17 |
| 27 | GPT-5.4 Nano | 9.25 | unverifiedT2 | 2026-03-17 |
| 28 | Qwen3.7-Plus | 9.14 | unverifiedT2 | 2026-06-02 |
| 29 | Inkling-Small | 8.29 | unverifiedT2 | 2026-07-30 |
| 30 | Kimi K2.6 | 8 | unverifiedT2 | 2026-04-20 |
| 31 | Grok 4.3 Beta | 8 | unverifiedT2 | 2026-04-17 |
| 32 | Gemini 3 Pro | 6.9 | unverifiedT2 | 2025-11-18 |
| 33 | Inkling | 5.43 | unverifiedT2 | 2026-07-15 |
| 34 | GPT-5.1 | 4.86 | unverifiedT2 | 2025-11-13 |
| 35 | GLM-5.1 | 4.57 | unverifiedT2 | 2026-04-07 |
| 36 | MiniMax-M3 | 3.71 | unverifiedT2 | 2026-06-01 |
| 37 | Kimi K2.5 | 3.14 | unverifiedT2 | 2026-02-02 |
| 38 | Nemotron 3 Ultra | 3.14 | unverifiedT2 | 2026-06-04 |
| 39 | Claude Sonnet 4.6 | 3.14 | unverifiedT2 | 2026-02-17 |
| 40 | Qwen 3.6 Plus | 2.86 | unverifiedT2 | 2026-04-01 |
| 41 | DeepSeek-V3.2-Exp | 2.86 | unverifiedT2 | 2025-09-29 |
| 42 | Gemini 2.5 Pro (Jun 2025) | 2 | unverifiedT2 | 2025-06-05 |
| 43 | GLM-4.7 | 1.71 | unverifiedT2 | 2025-12-22 |
| 44 | gpt-oss-20b | 1.43 | unverifiedT2 | 2025-08-05 |
| 45 | Gemma 4 31B IT | 1.43 | unverifiedT2 | 2026-04-02 |
| 46 | o3 | 1.4 | unverifiedT2 | 2024-12-20 |
| 47 | gpt-oss-120b | 1.14 | unverifiedT2 | 2025-08-05 |
| 48 | Gemini 3.1 Flash-Lite | 1.14 | unverifiedT2 | 2026-03-03 |
| 49 | Claude Sonnet 4.5 | 1.14 | unverifiedT2 | 2025-09-29 |
| 50 | GLM-4.6 | 1.14 | unverifiedT2 | 2025-09-30 |
| 51 | DeepSeek-R1 | 1.1 | unverifiedT2 | 2025-01-20 |
| 52 | Gemini 2.5 Flash (Jun 2025) | 1.1 | unverifiedT2 | 2025-06-17 |
| 53 | o4-mini | 0.6 | unverifiedT2 | 2025-04-16 |
| 54 | MiniMax-M2.7 | 0.57 | unverifiedT2 | 2026-03-18 |
| 55 | Claude Opus 4 | 0.3 | unverifiedT2 | 2025-05-22 |
| 56 | Qwen 3.6 35B-A3B | 0.29 | unverifiedT2 | 2026-04-14 |
| 57 | o3-mini | 0.29 | unverifiedT2 | 2025-01-31 |
| 58 | Claude Sonnet 4 | 0.29 | unverifiedT2 | 2025-05-22 |
| 59 | DeepSeek-V3 | 0 | unverifiedT2 | 2024-12-24 |
| 60 | Llama 3.1-8B | 0 | unverifiedT2 | 2024-07-23 |
| 61 | Llama 3.3 70B | 0 | unverifiedT2 | 2024-12-06 |
| 62 | Gemini 3.5 Flash-Lite | 0 | unverifiedT2 | 2026-07-21 |
| 63 | Gemma 3 27B | 0 | unverifiedT2 | 2025-03-12 |
| 64 | Claude 3.5 Haiku | 0 | unverifiedT2 | 2024-10-22 |
| 65 | Llama 4 Maverick | 0 | unverifiedT2 | 2025-04-05 |
| 66 | Llama 4 Scout | 0 | unverifiedT2 | 2025-04-05 |
| 67 | Claude Haiku 4.5 | 0 | unverifiedT2 | 2025-10-15 |
| 68 | Mistral Medium 3.5 | 0 | unverifiedT2 | 2026-04-29 |
| 69 | Mistral Small 3.1 | 0 | unverifiedT2 | 2025-03-17 |
| 70 | Qwen3-235B-A22B-Thinking (Jul 2025) | 0 | unverifiedT2 | 2025-07-25 |
| 71 | DeepSeek-V3 (Mar 2025) | 0 | unverifiedT2 | 2025-03-24 |
| 72 | GPT-4.1 mini | 0 | unverifiedT2 | 2025-04-14 |
| 73 | GPT-4.1 nano | 0 | unverifiedT2 | 2025-04-14 |
| 74 | GPT-4o (Nov 2024) | 0 | unverifiedT2 | 2024-05-13 |
| 75 | GPT-5 mini | 0 | unverifiedT2 | 2025-08-07 |
| 76 | GPT-5.5 Instant | 0 | unverifiedT2 | 2026-05-05 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
The benchmark's Benchmark Integrity Index is 94, corresponding to an integrity grade of A and indicating its health as a discriminator of frontier models under a disclosed harness. It ranks 23rd among 61 benchmarks in the benchmark universe. The weakest integrity component is discrimination, which reduces how much the benchmark's top-score distinctions can be trusted.
5 cited facts
The benchmark is not saturated: the top score is 32.3, leaving 67.7 points of headroom to the ceiling. Only two models cluster near the top, within a couple of points of each other. Because the field is not saturated, there is still genuine room to separate the top models, so small differences among leaders reflect real capability gaps.
5 cited facts
Scores for this benchmark are directly comparable because the evaluation uses a consistent harness, and the test set is held out rather than public, so results indicate task performance under the disclosed harness rather than deployed capability. The held-out status reduces the contamination scrutiny that would apply to a public test set; the contamination history reports construction from unpublished research and no known contamination. Grading uses a physics-informed automated pipeline customized for physics output formats, so scores are sensitive to harness choices such as code-execution tools and should be interpreted only under that exact disclosure.
5 cited facts
This benchmark measures frontier, research-level physics reasoning on novel unpublished problems, built by decomposing a small set of research challenges into checkpoints hand-crafted by active researchers from their own unpublished work. A human baseline is not documented. Its design assumes that expert-authored problems with verifiable answers proxy genuine frontier research reasoning, and that its auto-grader correctly checks complex physics outputs. The sharpest caveat is that near-floor scores limit discrimination among current models and auto-grading of symbolic or numeric physics is imperfect, so this benchmark is most likely to mislead when used to rank closely matched systems.
5 cited facts
CritPt (Complex Research using Integrated Thinking - Physics Test (CritPt)): CritPt (Complex Research using Integrated Thinking - Physics Test; reads as 'critical point') tests LLMs on unpublished, research-level physics reasoning across ~12 subfields, with machine-verifiable numeric/symbolic answers designed to resist guessing. Frontier models score in the single digits (best base model GPT-5 high approx 5.7%, approx 10% with coding tools).
GPT-5.6 Sol leads CritPt at 32.3. The full leaderboard above lists every recorded measurement, not just the headline number.
76 models have recorded scores on CritPt, spanning a score spread of 32.3.
tensor.news grades CritPt A for integrity (score 94/100), ranking #23 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — CritPt still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is held-out, which lowers — but does not eliminate — the risk that training data leaked into the questions.
Scores are reported under a consistent harness, so comparisons on CritPt are reasonably apples-to-apples.
Every CritPt measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites arxiv.org/abs/2509.26574.