Verdict
Llama 2-70B leads 8–0 across 10 shared benchmarks.
Llama 2-70B 8 · LLaMA-65B 0 · 2 tied · higher isn't always better — check the caveats below.
Reported scores — protocols may differ. How we compare
Commonsense Reasoning: Llama 2-70B leads 2–0 across 3 shared commonsense reasoning benchmarks (largest gap: Winogrande, 60.4 vs 54 — single evaluator).
Reasoning: Llama 2-70B leads 2–0 across 3 shared reasoning benchmarks (largest gap: ARC AI2, 71.07 vs 59.33 — single evaluator).
Llama 2-70B leads 8–0 across 10 shared benchmarks. Llama 2-70B leads 8 benchmarks and LLaMA-65B leads 0, with 2 tied. Higher isn't always better — see the integrity caveats on each benchmark.
Llama 2-70B and LLaMA-65B have 10 benchmarks in common in our data — those are the only rows where a direct, apples-to-apples comparison is drawn.
Llama 2-70B leads commonsense reasoning 2–0. Llama 2-70B leads reasoning 2–0.