Skip to content

Benchmark explainer

What is LMCA?

Crowdsourced LLM evaluation via blind pairwise human preference votes on real user conversations (Chatbot Arena).

LMCA as reported in Epoch AI's Capabilities Index CSV.

AINTEGRITY 100 / 100tensor.news
consistent harnesspublic test set

How it's scored

Metric
normalized performance in Epoch CSV, rendered as accuracy (%)
Score ceiling
100
Construction
Open conversational platform at lmarena.ai; users submit prompts, get two anonymous model responses side-by-side, vote for the better one; Elo ranks aggregate the votes.
Human baseline
Human preference vote IS the baseline (the metric being aggregated).

How trustworthy

Discrimination

100/100

Does it still separate models?

Saturation headroom

99.3/100

How far from ceiling / clustered at the top?

Contamination resistance

99.1/100

Public vs held-out; training-leak risk.

Harness comparability

100/100

Is it apples-to-apples, or mixed / vendor-optimized?

Freshness

100/100

How old is the benchmark?

Contamination history: User prompts and votes publicly released (lmsys/chatbot_arena_conversations) — prompts may end up in training data, inflating scores.

Limitations: Vote distribution skews to categories users prefer; new models lack sufficient vote depth (cold-start); user base demographics bias prompts; static benchmarks can saturate but Arena ranking continually shifts.

Who leads LMCA

ModelScoreEvidence
Claude Opus 574.47unverified
Claude Fable 570.99unverified
GPT-5.6 Sol69.66unverified
Claude Opus 4.867.66unverified
Claude Opus 4.665.6unverified

Compare the top LMCA scorers

Frequently asked questions

LMCA: LMCA as reported in Epoch AI's Capabilities Index CSV.

Claude Opus 5 leads LMCA at 74.47. The full leaderboard above lists every recorded measurement, not just the headline number.

111 models have recorded scores on LMCA, spanning a score spread of 71.2.

tensor.news grades LMCA A for integrity (score 100/100), ranking #3 of 62 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.

No — LMCA still has headroom and continues to discriminate between models rather than bunching them at the ceiling.

Its test set is unknown, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.

Scores are reported under a consistent harness, so comparisons on LMCA are reasonably apples-to-apples.

Every LMCA measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites arxiv.org/abs/2403.04132.

Follow the record