Verdict
GPT-6 Astra leads 12–2 across 16 shared benchmarks.
Claude Fable 5 2 · GPT-6 Astra 12 · 2 tied · higher isn't always better — check the caveats below.
Reported scores — protocols may differ. How we compare
Coding: Claude Fable 5 leads 2–1 across 3 shared coding benchmarks (largest gap: MirrorCode, 63.89 vs 46.7 — protocols undisclosed).
Reasoning: GPT-6 Astra leads 2–0 across 3 shared reasoning benchmarks (largest gap: GPQA diamond, 94.36 vs 81.14 — protocols undisclosed).
Unknown: GPT-6 Astra leads 3–0 across 3 shared unknown benchmarks (largest gap: EBR-bench, 76.19 vs 39.52 — protocols undisclosed).
Math: GPT-6 Astra leads 1–0 across 2 shared math benchmarks (largest gap: FrontierMath-Tier-4-v2-Private, 97.6 vs 90.2 — protocols undisclosed).
Vendor claims — compare against the measured scores above. A claim is what a developer says about its own model, not an independent measurement.
Agentic / Tool Use
“The longer and more complex the task, the larger Fable 5's lead over our other models”
Coding
“It is state-of-the-art on nearly all tested benchmarks of AI capability, showing exceptional performance in software engineering, knowledge work, vision, scientific research, and many other areas”
General
“Fable 5's capabilities exceed those of any model we've ever made generally available”
Multimodal
“It can extract precise numbers from detailed scientific figures and can perform complex vision-based tasks like rebuilding a web app's source code from screenshots alone”
No vendor claims on file for GPT-6 Astra.
GPT-6 Astra leads 12–2 across 16 shared benchmarks. Claude Fable 5 leads 2 benchmarks and GPT-6 Astra leads 12, with 2 tied. Higher isn't always better — see the integrity caveats on each benchmark.
Claude Fable 5 and GPT-6 Astra have 16 benchmarks in common in our data — those are the only rows where a direct, apples-to-apples comparison is drawn.
GPT-6 Astra leads unknown 3–0. GPT-6 Astra leads reasoning 2–0.