Benchmark explainer
What is LMCA?
Crowdsourced LLM evaluation via blind pairwise human preference votes on real user conversations (Chatbot Arena).
LMCA as reported in Epoch AI's Capabilities Index CSV.
How it's scored
- Metric
- normalized performance in Epoch CSV, rendered as accuracy (%)
- Score ceiling
- 100
- Construction
- Open conversational platform at lmarena.ai; users submit prompts, get two anonymous model responses side-by-side, vote for the better one; Elo ranks aggregate the votes.
- Human baseline
- Human preference vote IS the baseline (the metric being aggregated).
How trustworthy
Discrimination
100/100Does it still separate models?
Saturation headroom
99.3/100How far from ceiling / clustered at the top?
Contamination resistance
99.1/100Public vs held-out; training-leak risk.
Harness comparability
100/100Is it apples-to-apples, or mixed / vendor-optimized?
Freshness
100/100How old is the benchmark?
Contamination history: User prompts and votes publicly released (lmsys/chatbot_arena_conversations) — prompts may end up in training data, inflating scores.
Limitations: Vote distribution skews to categories users prefer; new models lack sufficient vote depth (cold-start); user base demographics bias prompts; static benchmarks can saturate but Arena ranking continually shifts.
Who leads LMCA
| Model | Score | Evidence |
|---|---|---|
| Claude Opus 5 | 74.47 | unverified |
| Claude Fable 5 | 70.99 | unverified |
| GPT-5.6 Sol | 69.66 | unverified |
| Claude Opus 4.8 | 67.66 | unverified |
| Claude Opus 4.6 | 65.6 | unverified |
Compare the top LMCA scorers
Frequently asked questions
LMCA: LMCA as reported in Epoch AI's Capabilities Index CSV.
Claude Opus 5 leads LMCA at 74.47. The full leaderboard above lists every recorded measurement, not just the headline number.
111 models have recorded scores on LMCA, spanning a score spread of 71.2.
tensor.news grades LMCA A for integrity (score 100/100), ranking #3 of 62 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — LMCA still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is unknown, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Scores are reported under a consistent harness, so comparisons on LMCA are reasonably apples-to-apples.
Every LMCA measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites arxiv.org/abs/2403.04132.