Skip to content

Gemini 3.1 Pro vs Qwen 3.6 Max (Preview)

Verdict

Gemini 3.1 Pro leads 5–1 across 6 shared benchmarks.

Gemini 3.1 Pro 5 · Qwen 3.6 Max (Preview) 1 · higher isn't always better — see caveats.

Per-benchmark head-to-head

Reported scores — protocols may differ. How we compare

Sort
Gemini 3.1 ProQwen 3.6 Max (Preview)

Capabilities

Reasoning

Reasoning: Gemini 3.1 Pro leads 3–0 across 3 shared reasoning benchmarks (largest gap: SimpleQA Verified, 77.3 vs 56.93 — single evaluator).

4 reproduced2 unverified

What the labs claim

Vendor claims — compare against the measured scores above. A claim is what a developer says about its own model, not an independent measurement.

Gemini 3.1 Pro

Agentic / Tool Use

  • Gemini 3.1 Pro was evaluated across a range of benchmarks, including reasoning, multimodal capabilities, agentic tool use, multi-lingual performance, and long-context.

Multimodal

  • a suite of highly capable, natively multimodal reasoning models

Reasoning

  • Gemini 3.1 Pro is Google's most advanced model for complex tasks.

Qwen 3.6 Max (Preview)

Agentic / Tool Use

  • Compared to Qwen3.6-Plus, this preview release brings stronger world knowledge and instruction following, along with significant agentic coding improvements across a wide range of benchmarks.
  • improved real-world agent and knowledge reliability performance

Coding

  • It achieves the top score on six major coding benchmarks - SWE-bench Pro, Terminal-Bench 2.0, SkillsBench, QwenClawBench, QwenWebBench, and SciCode - with substantial gains over its predecessor
A higher number is not always a better model. Each score is task performance under a disclosed harness. Of the 12 measurements across 6 shared benchmarks, 12 single-evaluator or undisclosed-protocol reports; 83% are independently reproduced. Qwen 3.6 Max (Preview) pricing is cross-check disputed — weigh cost claims with the same caution as benchmark wins. Follow any benchmark to its integrity grade before reading a win as decisive.

Follow the record