Benchmark explainer
What is LiveCodeBench?
Competitive programming on problems released after typical training cutoffs
LiveCodeBench — a continuously updated code-generation benchmark pulling fresh problems from Codeforces/LeetCode/AtCoder after typical training cutoffs, scored as pass@1 (%); designed to stay contamination-resistant as the collection window advances.
How it's scored
- Metric
- pass@1
- Score ceiling
- 100
How trustworthy
Discrimination
100/100Does it still separate models?
Saturation headroom
91.8/100How far from ceiling / clustered at the top?
Contamination resistance
100/100Public vs held-out; training-leak risk.
Harness comparability
100/100Is it apples-to-apples, or mixed / vendor-optimized?
Freshness
100/100How old is the benchmark?
Limitations: Problems are public once collected; the effective cutoff depends on the release window the evaluator chose, which papers don't always disclose
Who leads LiveCodeBench
| Model | Score | Evidence |
|---|---|---|
| DeepSeek-R1 | 65.9 | self-reported |
| DeepSeek-R1-Distill-Llama-70B | 57.5 | self-reported |
| DeepSeek-R1-Distill-Qwen-32B | 57.2 | self-reported |
| DeepSeek-R1-Distill-Qwen-14B | 53.1 | self-reported |
| DeepSeek-R1-Distill-Llama-8B | 39.6 | self-reported |
Frequently asked questions
LiveCodeBench (LiveCodeBench (rolling fresh-problem code benchmark)): LiveCodeBench — a continuously updated code-generation benchmark pulling fresh problems from Codeforces/LeetCode/AtCoder after typical training cutoffs, scored as pass@1 (%); designed to stay contamination-resistant as the collection window advances.
DeepSeek-R1 leads LiveCodeBench at 65.9 — self-reported by the lab. The full leaderboard above lists every recorded measurement, not just the headline number.
8 models have recorded scores on LiveCodeBench, spanning a score spread of 49.
tensor.news grades LiveCodeBench A for integrity (score 98/100), ranking #16 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — LiveCodeBench still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is public, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Scores are reported under a consistent harness, so comparisons on LiveCodeBench are reasonably apples-to-apples.
Every LiveCodeBench measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites arxiv.org/abs/2403.04132.