Crowdsourced LLM evaluation via blind pairwise human preference votes on real user conversations (Chatbot Arena).
Strongest on discrimination, weakest on contamination resistance. Caveat: the public test set does not rule out training-data contamination.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | Claude Opus 5 | 74.47 | unverifiedT1 | 2026-07-24 |
| 2 | Claude Fable 5 | 70.99 | unverifiedT1 | 2026-06-09 |
| 3 | GPT-5.6 Sol | 69.66 | unverifiedT1 | 2026-07-09 |
| 4 | Claude Opus 4.8 | 67.66 | unverifiedT1 | 2026-05-28 |
| 5 | Claude Opus 4.6 | 65.6 | unverifiedT1 | 2026-02-05 |
| 6 | GPT-5.5 | 63.85 | unverifiedT1 | 2026-04-23 |
| 7 | GPT-5.5 Pro | 63.46 | unverifiedT1 | 2026-04-23 |
| 8 | Gemini 3.1 Pro | 63.32 | unverifiedT1 | 2026-02-19 |
| 9 | Kimi K3 | 62.05 | unverifiedT1 | 2026-07-16 |
| 10 | GPT-5.6 Terra | 61.45 | unverifiedT1 | 2026-07-09 |
| 11 | Claude Opus 4.7 | 61.41 | unverifiedT1 | 2026-04-16 |
| 12 | GPT-5.4 | 61.14 | unverifiedT1 | 2026-03-05 |
| 13 | Gemini 3.7 Flash | 59.27 | unverifiedT1 | 2026-08-13 |
| 14 | Muse Spark 1.1 | 58.72 | unverifiedT1 | 2026-07-09 |
| 15 | Claude Sonnet 5 | 58.01 | unverifiedT1 | 2026-06-30 |
| 16 | GPT-5.6 Luna | 57.11 | unverifiedT1 | 2026-07-09 |
| 17 | Grok 4.6 | 57 | unverifiedT1 | 2026-08-12 |
| 18 | Muse Spark 1.2 | 56.88 | unverifiedT1 | 2026-08-05 |
| 19 | Gemini 3.5 Flash | 55.47 | unverifiedT1 | 2026-05-19 |
| 20 | Claude Sonnet 4.6 | 54.72 | unverifiedT1 | 2026-02-17 |
| 21 | Qwen 3.8 Max | 54.4 | unverifiedT1 | 2026-08-02 |
| 22 | GLM-5.2 | 53.91 | unverifiedT1 | 2026-06-16 |
| 23 | Grok 4.5 | 53.18 | unverifiedT1 | 2026-07-08 |
| 24 | Gemini 3.6 Flash | 52.86 | unverifiedT1 | 2026-07-21 |
| 25 | Claude Opus 4.5 | 52.31 | unverifiedT1 | 2025-11-24 |
| 26 | Qwen3.7-Max | 51.79 | unverifiedT1 | 2026-05-19 |
| 27 | GPT-5.2 | 51.66 | unverifiedT1 | 2025-12-11 |
| 28 | GPT-5.1 | 51.64 | unverifiedT1 | 2025-11-13 |
| 29 | Gemini 3 Flash | 50.74 | unverifiedT1 | 2025-12-17 |
| 30 | Qwen 3.6 Max (Preview) | 50.04 | unverifiedT1 | 2026-04-20 |
| 31 | DeepSeek V4 Flash 0731 | 49.09 | unverifiedT1 | 2026-07-31 |
| 32 | DeepSeek-V4-Pro | 48.47 | unverifiedT1 | 2026-04-24 |
| 33 | GPT-5.4 Mini | 47.99 | unverifiedT1 | 2026-03-17 |
| 34 | GPT-5 | 47.05 | unverifiedT1 | 2025-08-07 |
| 35 | o3 | 46.71 | unverifiedT1 | 2025-04-16 |
| 36 | Gemma 4 31B IT | 46.18 | unverifiedT1 | 2026-04-02 |
| 37 | Claude Sonnet 4.5 | 45.6 | unverifiedT1 | 2025-09-29 |
| 38 | Grok 4.20 | 45.48 | unverifiedT1 | 2026-02-17 |
| 39 | o3-pro | 45.27 | unverifiedT1 | 2025-06-10 |
| 40 | Grok 4.3 Beta | 45.09 | unverifiedT1 | 2026-04-17 |
| 41 | Qwen3.5 397B-A17B | 44.58 | unverifiedT1 | 2026-02-13 |
| 42 | Gemini 3.5 Flash-Lite | 44.24 | unverifiedT1 | 2026-07-21 |
| 43 | Inkling | 44.21 | unverifiedT1 | 2026-07-15 |
| 44 | Qwen3.7-Plus | 44.21 | unverifiedT1 | 2026-06-02 |
| 45 | Claude Opus 4 | 44.04 | unverifiedT1 | 2025-05-22 |
| 46 | Kimi K2.6 | 43.84 | unverifiedT1 | 2026-04-20 |
| 47 | Claude Opus 4.1 | 43.62 | unverifiedT1 | 2025-08-05 |
| 48 | Nemotron 3 Ultra | 43.41 | unverifiedT1 | 2026-06-04 |
| 49 | GPT-5.4 Nano | 43.36 | unverifiedT1 | 2026-03-17 |
| 50 | Qwen 3.5 Plus (hosted 397B-A17B) | 42.78 | unverifiedT1 | 2026-02-16 |
| 51 | DeepSeek-V4-Flash | 42.21 | unverifiedT1 | 2026-04-24 |
| 52 | Gemini 3.1 Flash-Lite | 41.19 | unverifiedT1 | 2026-03-03 |
| 53 | Gemini 2.5 Pro (Jun 2025) | 40.91 | unverifiedT1 | 2025-06-05 |
| 54 | Qwen3.6 27B | 40.62 | unverifiedT1 | 2026-04-22 |
| 55 | GPT-5 mini | 40.24 | unverifiedT1 | 2025-08-07 |
| 56 | MiniMax-M3 | 39.65 | unverifiedT1 | 2026-06-01 |
| 57 | Qwen 3.6 Plus | 38.93 | unverifiedT1 | 2026-03-31 |
| 58 | Qwen 3.6 Flash | 36.45 | unverifiedT1 | 2026-04-27 |
| 59 | Claude Haiku 4.5 | 36.4 | unverifiedT1 | 2025-10-15 |
| 60 | Qwen 3.6 35B-A3B | 34.94 | unverifiedT1 | 2026-04-14 |
| 61 | Gemma 4 26B A4B | 34.92 | unverifiedT1 | 2026-04-02 |
| 62 | Qwen3.5-35B-A3B | 34.67 | unverifiedT1 | 2026-02-24 |
| 63 | Qwen3-235B-A22B-Thinking (Jul 2025) | 34.48 | unverifiedT1 | 2025-07-25 |
| 64 | Qwen 3.5 Flash (hosted 35B-A3B) | 34.28 | unverifiedT1 | 2026-02-25 |
| 65 | DeepSeek-V3.2-Exp | 34.25 | unverifiedT1 | 2025-09-29 |
| 66 | Claude Sonnet 4 | 34.12 | unverifiedT1 | 2025-05-22 |
| 67 | DeepSeek-V3.2 | 33.91 | unverifiedT1 | 2025-12-01 |
| 68 | Qwen3-Max | 33.26 | unverifiedT1 | 2025-09-24 |
| 69 | Gemini 2.5 Flash (Jun 2025) | 32.34 | unverifiedT1 | 2025-06-17 |
| 70 | o4-mini | 31.18 | unverifiedT1 | 2025-04-16 |
| 71 | Mistral Medium 3.5 | 30.66 | unverifiedT1 | 2026-04-28 |
| 72 | GPT-4.1 | 30.13 | unverifiedT1 | 2025-04-14 |
| 73 | Qwen3-235B-A22B | 29.41 | unverifiedT1 | 2025-04-28 |
| 74 | Qwen3.5-9B | 28.79 | unverifiedT1 | 2026-02-24 |
| 75 | DeepSeek-V3.1 | 28.58 | unverifiedT1 | 2025-08-21 |
| 76 | Qwen3-235B-A22B-Instruct (Jul 2025) | 27.67 | unverifiedT1 | 2025-07-25 |
| 77 | Qwen3-30B-A3B-Instruct (Jul 2025) | 26.32 | unverifiedT1 | 2025-07-29 |
| 78 | o1 | 26.2 | unverifiedT1 | 2024-12-17 |
| 79 | gpt-oss-120b | 26.05 | unverifiedT1 | 2025-08-05 |
| 80 | GPT-4.1 mini | 24.85 | unverifiedT1 | 2025-04-14 |
| 81 | Qwen3-30B-A3B-Thinking (Jul 2025) | 22.98 | unverifiedT1 | 2025-07-30 |
| 82 | o3-mini | 22.32 | unverifiedT1 | 2025-01-31 |
| 83 | Qwen3-14B | 21.4 | unverifiedT1 | 2025-04-29 |
| 84 | Gemini 2.5 Flash-Lite (Jun 2025) | 21.26 | unverifiedT1 | 2025-06-17 |
| 85 | Llama 3.3 70B | 20.61 | unverifiedT1 | 2024-12-06 |
| 86 | Qwen3-32B | 20.33 | unverifiedT1 | 2025-04-28 |
| 87 | GPT-4 (Jun 2023) | 20.07 | unverifiedT1 | 2023-06-13 |
| 88 | Claude 3 Opus | 20 | unverifiedT1 | 2024-02-29 |
| 89 | GPT-4o (May 2024) | 19.52 | unverifiedT1 | 2024-05-13 |
| 90 | Llama 4 Maverick | 18.67 | unverifiedT1 | 2025-04-05 |
| 91 | Qwen3-30B-A3B | 18.61 | unverifiedT1 | 2025-04-28 |
| 92 | DeepSeek-V3 (Mar 2025) | 18.2 | unverifiedT1 | 2025-03-24 |
| 93 | Llama 3.1-70B | 17.47 | unverifiedT1 | 2024-07-23 |
| 94 | gpt-oss-20b | 17.07 | unverifiedT1 | 2025-08-05 |
| 95 | Qwen2.5-72B | 15.79 | unverifiedT1 | 2024-09-19 |
| 96 | Gemma 3 27B | 14.45 | unverifiedT1 | 2025-03-12 |
| 97 | Llama 4 Scout | 14.07 | unverifiedT1 | 2025-04-05 |
| 98 | GPT-4o mini | 12.2 | unverifiedT1 | 2024-07-18 |
| 99 | GPT-4 Turbo (Nov 2023) | 11.51 | unverifiedT1 | 2023-11-06 |
| 100 | GPT-3.5 Turbo (Jan 2024) | 11.35 | unverifiedT1 | 2024-01-25 |
| 101 | Claude 3 Haiku | 10.4 | unverifiedT1 | 2024-03-07 |
| 102 | Qwen3-8B | 10.38 | unverifiedT1 | 2025-04-28 |
| 103 | GPT-5 nano | 9.33 | unverifiedT1 | 2025-08-07 |
| 104 | Gemma 2 27B | 8.35 | unverifiedT1 | 2024-06-24 |
| 105 | Qwen2.5-7B | 7.55 | unverifiedT1 | 2024-09-19 |
| 106 | Ministral 3B | 6.45 | unverifiedT1 | 2024-10-16 |
| 107 | GPT-4.1 nano | 6.42 | unverifiedT1 | 2025-04-14 |
| 108 | Llama 3.1-8B | 6.32 | unverifiedT1 | 2024-07-23 |
| 109 | Command R+ | 5.89 | unverifiedT1 | 2024-08-30 |
| 110 | Gemma 3 12B | 5.25 | unverifiedT1 | 2025-03-12 |
| 111 | Gemma 3 4B | 3.27 | unverifiedT1 | 2025-03-12 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
This benchmark scores 111 models, with a spread of 71.2 points between the best and worst performers. A spread that wide indicates substantial separation between models, so rank differences at the top and bottom of the table are likely to reflect real capability gaps rather than noise.
3 cited facts
This benchmark's integrity score is 100, corresponding to grade A, placing it 3rd among 62 benchmarks in the universe. The weakest component of its integrity breakdown is contamination, which narrows how much confidence readers should place in top scores since test-set privacy is unknown. Read as the benchmark's health as a discriminator of frontier models under a disclosed harness, this score speaks to the evaluation instrument rather than to any model's capability.
6 cited facts
The benchmark is not saturated: its top score is 74.47, leaving 25.53 points of headroom to the ceiling, and only 1 model sits within a couple of points of the top. Because it is not saturated, there is still genuine room to separate the field at the top, so differences among the leading models can be read as meaningful rather than noise.
5 cited facts
Scores from this benchmark are directly comparable because the harness is consistent, but they should not be compared head-to-head with scores from a different harness due to low cross-harness comparability. The test set privacy is unknown, so it is indeterminate whether high scores deserve the extra contamination scrutiny that a public test set would warrant, whereas a held-out test set would lessen that concern. The contamination history notes that user prompts and votes were publicly released, which may allow prompts to enter training data and inflate scores, and the harness notes indicate that optimized-scaffold flags are carried forward, underscoring that reported scores reflect task performance under a disclosed harness rather than deployed capability.
4 cited facts
This benchmark measures crowdsourced evaluation of large language models through blind pairwise human preference votes on real user conversations, with the human preference vote itself serving as the baseline metric that is aggregated. The test is constructed on an open conversational platform where users submit prompts and receive a pair of anonymous model responses side by side, then vote for the better one; a rating system aggregates these votes into rankings. Its design assumes that real-user prompts evaluated through blind pairwise preference provide the most production-relevant quality signal, and that the rating system captures incremental capability differences cleanly. The sharpest caveat is that the vote distribution skews toward categories users prefer and that new models lack sufficient vote depth, so the ranking may mislead about a model's true capability early on.
5 cited facts
LMCA: LMCA as reported in Epoch AI's Capabilities Index CSV.
Claude Opus 5 leads LMCA at 74.47. The full leaderboard above lists every recorded measurement, not just the headline number.
111 models have recorded scores on LMCA, spanning a score spread of 71.2.
tensor.news grades LMCA A for integrity (score 100/100), ranking #3 of 62 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — LMCA still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is unknown, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Scores are reported under a consistent harness, so comparisons on LMCA are reasonably apples-to-apples.
Every LMCA measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites arxiv.org/abs/2403.04132.