Skip to content

Only 29% of AI model performance claims are independently verified

tensor.news desk · 2026-07-24

Of 1,943 performance claims in the tensor.news record, just 573 have been independently reproduced. The remaining 1,370 scores are self-reported or unverified, meaning the evidence base for frontier models remains largely unconfirmed as of July 2026.

Who leads and on what evidence

Anthropic's Claude Fable 5 holds the most benchmark leads — six in total — including CursorBench, FrontierMath-Tier-4-v2-Private, GBAEval, Remote Labor Index, SimpleBench, and WeirdML. All claims for Claude Fable 5 are independently reproduced, and its best recorded score is 99.72.

OpenAI's GPT-5.5 Pro leads four benchmarks: Chess Puzzles, CritPt, FrontierMath-Tiers-1-3-v2-Private, and OTIS Mock AIME 2024–2025, with a best score of 100. GPT-5.5, a distinct model, also leads ARC-AGI-2, CL-bench Life, ExploitBench, and the same OTIS Mock AIME, with a perfect score. Every benchmark lead held by these OpenAI models is independently reproduced.

Google DeepMind's Gemini 3.1 Pro leads ARC-AGI, HLE, and SimpleQA Verified, all independently reproduced, with a top score of 98. Other models in the top 10 — GPT-5, Gemini 3 Pro, Claude Opus 4.7, and even GPT-4 (Mar 2023) and Llama 3.1-405B — also carry the same independently reproduced status for their leads. No frontier model in the current top 10 has a disputed or unverified lead claim.

Where benchmarks are saturating or eroding

Four benchmarks receive a perfect integrity score of 100: HLE, SimpleQA Verified, Chess Puzzles, and APEX-Agents. ARC-AGI-2, at 99, rounds out the top five, all graded A. In each case, the weakest component is saturation — scores on these tests have climbed high enough that differentiation among top models is becoming compressed.

At the other end, five benchmarks score between 54 and 62, graded D or C, with contamination as the weakest component. LAMBADA (54), ANLI (59), OpenBookQA (61), PIQA (62), and Winogrande (62) are all eroding under contamination pressure. These legacy benchmarks no longer provide reliable signals of model capability, and their continued use risks overstating progress.

What the money says

Among the 68 publicly priced models in the tensor.news pricing index, only 29 prices are corroborated across multiple sources. Thirty are disputed, and 9 rely on a single source. More listed prices are contested than confirmed.

The cheapest frontier-adjacent models start with GPT-5 nano, whose blended price of $0.138 per million tokens is disputed. Mistral NeMo ($0.15) and Qwen3-235B-A22B ($0.30) are also disputed. Two of the five cheapest — GPT-4.1 nano ($0.175) and GPT-4o mini ($0.262) — are not disputed.

The most expensive models concentrate at extreme price points. GPT-5.5 Pro and GPT-5.4 Pro share a blended price of $67.50, both corroborated. GPT-5.2 Pro ($57.75) is disputed. GPT-5 Pro ($41.25) and o3-pro ($35.00) round out the top tier and are corroborated. The spread from $0.138 to $67.50 per million tokens is a 489-fold cost range for accessing a frontier model API.

What to watch

  • GLM-5.2, with scores recorded on June 16, 2026, has four of its ten listed benchmarks unverified: WeirdML (70.12), CritPt (20.86), ARC-AGI (77), and ARC-AGI-2 (22.78). Independent reproduction of those numbers would clarify the model's true standing.
  • HLE, SimpleQA Verified, and ARC-AGI-2 are all flagged for saturation, suggesting that the next generation of models will require new, harder evaluation suites to discriminate performance beyond current ceilings.
  • With 30 disputed prices versus 29 corroborated, the pricing data for frontier models remains unreliable. Resolution of the disputes, especially for models like GPT-5.2 Pro and GPT-5 nano, will affect cost calculations for builders and enterprise buyers.