CL-bench: A Benchmark for Context Learning
Integrity rank #38 of 61 · 19 models scored · top score 27.9 · GPT-5.4
Context learning - acquiring and applying novel knowledge presented only in-context
Strongest on contamination resistance, weakest on discrimination. Caveat: the public test set does not rule out training-data contamination.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | GPT-5.4 | 27.9 | unverifiedT2 | 2026-03-05 |
| 2 | GPT-5.1 | 23.7 | unverifiedT2 | 2025-11-13 |
| 3 | Grok 4.20 | 22.2 | unverifiedT2 | 2026-02-17 |
| 4 | Claude Opus 4.5 | 21.1 | unverifiedT2 | 2025-11-24 |
| 5 | Gemini 3.1 Pro | 20.8 | unverifiedT2 | 2026-02-19 |
| 6 | Claude Opus 4.6 | 20.7 | unverifiedT2 | 2026-02-05 |
| 7 | Qwen 3.6 Plus | 20.3 | unverifiedT2 | 2026-04-01 |
| 8 | Qwen 3.5 Plus (hosted 397B-A17B) | 19.8 | unverifiedT2 | 2026-02-16 |
| 9 | Kimi K2.5 | 19.3 | unverifiedT2 | 2026-02-02 |
| 10 | GLM-5 | 18.7 | unverifiedT2 | 2026-02-11 |
| 11 | GPT-5.2 | 18.2 | unverifiedT2 | 2025-12-11 |
| 12 | o3 | 17.8 | unverifiedT2 | 2024-12-20 |
| 13 | Kimi K2 Thinking | 17.6 | unverifiedT2 | 2025-11-06 |
| 14 | GLM-4.7 | 15.9 | unverifiedT2 | 2025-12-22 |
| 15 | Gemini 3 Pro | 15.8 | unverifiedT2 | 2025-11-18 |
| 16 | Qwen3-Max | 14.5 | unverifiedT2 | 2025-09-05 |
| 17 | DeepSeek-V3.2-Exp | 13.2 | unverifiedT2 | 2025-09-29 |
| 18 | DeepSeek-V3.2 | 12.4 | unverifiedT2 | 2025-12-01 |
| 19 | MiniMax-M2.5 | 11.4 | unverifiedT2 | 2026-02-12 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
Mid-table integrity — 82, grade B, ranked 29 of 51. Solid enough to take results seriously, not clean enough to settle close calls; corroborate tight margins on a higher-ranked benchmark. This score represents the benchmark's health as a discriminator of frontier models under a disclosed harness, though age complicates the assessment.
4 cited facts
Scores have not run out of room here, which is what keeps the top of the table meaningful — leading gaps are measured differences, not artifacts of a maxed-out scale. Scores are clustered near the top, indicating they are essentially within a couple of points of each other. The top score is 27.9 and the score headroom is 72.1, leaving considerable room before the ceiling is reached.
4 cited facts
Trust the ordering more than the level: a consistent harness makes the ranking meaningful, while the absolute scores remain artifacts of this particular setup. Because the test set is public, unusually high scores warrant closer inspection for potential contamination. The deliberate inclusion of knowledge absent from pre‑training and rubric‑based, model‑assisted grading makes this benchmark particularly sensitive to LLM‑judge behavior and contamination concerns挣钱
4 cited facts
CL-bench (CL-bench: A Benchmark for Context Learning): CL-bench (A Benchmark for Context Learning) tests whether models can absorb and apply genuinely novel in-context knowledge (fictional legal systems, invented game rules, novel financial instruments, empirically derived laws) rather than recall pretrained facts. Average solving rate approx 17.2%; best model GPT-5.1 approx 23.7%.
GPT-5.4 leads CL-bench at 27.9. The full leaderboard above lists every recorded measurement, not just the headline number.
19 models have recorded scores on CL-bench, spanning a score spread of 16.5.
tensor.news grades CL-bench B for integrity (score 82/100), ranking #38 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — CL-bench still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is public, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Scores are reported under a consistent harness, so comparisons on CL-bench are reasonably apples-to-apples.
Every CL-bench measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites arxiv.org/abs/2602.03587.