Verdict
Claude Fable 5 leads 4–0 across 4 shared benchmarks.
Claude Fable 5 4 · Llama 3.1-405B 0 · higher isn't always better — check the caveats below.
Reported scores — protocols may differ. How we compare
Reasoning: Claude Fable 5 leads 2–0 across 2 shared reasoning benchmarks (largest gap: SimpleBench, 78.28 vs 7.6 — single evaluator).
Vendor claims — compare against the measured scores above. A claim is what a developer says about its own model, not an independent measurement.
Agentic / Tool Use
“The longer and more complex the task, the larger Fable 5's lead over our other models”
Coding
“It is state-of-the-art on nearly all tested benchmarks of AI capability, showing exceptional performance in software engineering, knowledge work, vision, scientific research, and many other areas”
General
“Fable 5's capabilities exceed those of any model we've ever made generally available”
Multimodal
“It can extract precise numbers from detailed scientific figures and can perform complex vision-based tasks like rebuilding a web app's source code from screenshots alone”
Context Handling
“These are multilingual and have a significantly longer context length of 128K, state-of-the-art tool use, and overall stronger reasoning capabilities.”
General
“Llama 3.1 405B is in a class of its own, with unmatched flexibility, control, and state-of-the-art capabilities that rival the best closed source models.”
“Llama 3.1 405B is the first openly available model that rivals the top AI models when it comes to state-of-the-art capabilities in general knowledge, steerability, math, tool use”
Claude Fable 5 leads 4–0 across 4 shared benchmarks. Claude Fable 5 leads 4 benchmarks and Llama 3.1-405B leads 0. Higher isn't always better — see the integrity caveats on each benchmark.
Claude Fable 5 and Llama 3.1-405B have 4 benchmarks in common in our data — those are the only rows where a direct, apples-to-apples comparison is drawn.
Claude Fable 5 leads reasoning 2–0.