LLM Creative Story-Writing Benchmark
Integrity rank #45 of 61 · 41 models scored · top score 86 · GPT-5
Creative short-fiction quality with hard constraint satisfaction
Strongest on harness comparability, weakest on contamination resistance. Caveat: the public test set does not rule out training-data contamination.
Reported scores — protocols may differ. How we compare
| # | Model | Score | Evidence | Measured |
|---|---|---|---|---|
| 1 | GPT-5 | 86 | unverifiedT1 | 2025-08-07 |
| 2 | Kimi K2 (Jul 2025) | 85.6 | unverifiedT1 | 2025-07-11 |
| 3 | Claude Opus 4.1 | 84.7 | unverifiedT1 | 2025-08-05 |
| 4 | o3-pro | 84.4 | unverifiedT1 | 2025-06-10 |
| 5 | o3 | 83.9 | unverifiedT1 | 2024-12-20 |
| 6 | Gemini 2.5 Pro (Jun 2025) | 83.8 | unverifiedT1 | 2025-06-05 |
| 7 | Claude Opus 4 | 83.6 | unverifiedT1 | 2025-05-22 |
| 8 | GPT-5 mini | 83.1 | unverifiedT1 | 2025-08-07 |
| 9 | DeepSeek-R1 | 83 | unverifiedT1 | 2025-01-20 |
| 10 | Qwen3-235B-A22B | 83 | unverifiedT1 | 2025-04-28 |
| 11 | Qwen3-235B-A22B-Thinking (Jul 2025) | 82.4 | unverifiedT1 | 2025-07-25 |
| 12 | DeepSeek-R1 (May 2025) | 81.9 | unverifiedT1 | 2025-05-28 |
| 13 | GPT-4o (Nov 2024) | 81.8 | unverifiedT1 | 2024-05-13 |
| 14 | Claude Sonnet 4 | 81.4 | unverifiedT1 | 2025-05-22 |
| 15 | Claude 3.7 Sonnet | 81.1 | unverifiedT1 | 2025-02-24 |
| 16 | Gemini 2.5 Pro (May 2025) | 80.9 | unverifiedT1 | 2025-05-06 |
| 17 | Gemini 2.5 Pro (Mar 2025) | 80.5 | unverifiedT1 | 2025-03-25 |
| 18 | Claude 3.5 Sonnet (October 2024) | 80.3 | unverifiedT1 | 2024-10-22 |
| 19 | Gemma 3 27B | 79.9 | unverifiedT1 | 2025-03-12 |
| 20 | Mistral Medium 3 | 77.3 | unverifiedT1 | 2025-05-07 |
| 21 | gpt-oss-120b | 77.1 | unverifiedT1 | 2025-08-05 |
| 22 | DeepSeek-V3 (Mar 2025) | 77 | unverifiedT1 | 2025-03-24 |
| 23 | Grok 4 | 76.9 | unverifiedT1 | 2025-07-09 |
| 24 | Gemini 2.5 Flash (Apr 2025) | 76.5 | unverifiedT1 | 2025-04-17 |
| 25 | Grok 3 | 76.4 | unverifiedT1 | 2025-02-17 |
| 26 | GPT-4.5 | 75.6 | unverifiedT1 | 2025-02-27 |
| 27 | o4-mini | 75 | unverifiedT1 | 2025-04-16 |
| 28 | Gemini 2.0 Flash Thinking (Jan 2025) | 73.8 | unverifiedT1 | 2025-01-21 |
| 29 | Claude 3.5 Haiku | 73.5 | unverifiedT1 | 2024-10-22 |
| 30 | Grok-3 mini | 73.5 | unverifiedT1 | 2025-02-19 |
| 31 | Qwen2.5-Max | 72.9 | unverifiedT1 | 2025-01-25 |
| 32 | Gemini 2.0 Flash (Dec 2024) | 71.5 | unverifiedT1 | 2024-12-11 |
| 33 | o1 | 70.2 | unverifiedT1 | 2024-12-05 |
| 34 | Mistral Large 2 (Jul 2024) | 69 | unverifiedT1 | 2024-07-24 |
| 35 | GPT-4o mini | 67.2 | unverifiedT1 | 2024-07-18 |
| 36 | o1-mini | 64.9 | unverifiedT1 | 2024-09-12 |
| 37 | Grok-2 (Dec 2024) | 63.6 | unverifiedT1 | 2024-08-13 |
| 38 | Phi-4 | 62.6 | unverifiedT1 | 2024-12-12 |
| 39 | Llama 4 Maverick | 62 | unverifiedT1 | 2025-04-05 |
| 40 | o3-mini | 61.7 | unverifiedT1 | 2025-01-31 |
| 41 | Amazon Nova Pro | 60.5 | unverifiedT1 | 2024-12-03 |
Continue from this leaderboard into the tightest head-to-head reads.
Weighted composite. Each dimension opens to the source-backed sub-signals it is computed from.
Claims drawn from cited facts, not live model generation.
This benchmark scores 41 models with a spread of 25.5 points between best and worst. Such a wide spread suggests that differences in model rankings likely reflect genuine capability gaps.
2 cited facts
The benchmark's Integrity Index of 76 (grade B) indicates its health as a discriminator of frontier models, ranking 39th out of 53 benchmarks. Its weakest component is contamination, which narrows how much to trust the highest scores.
4 cited facts
Public test data undercuts part of what the consistent harness buys: rankings stay clean, but a high score carries an open question about training-set exposure that a held-out set would have closed. High comparability allows direct head-to-head comparisons, while the public test set requires careful scrutiny for possible contamination. Although all prompts and reference stories are public, the large combinatorial element space limits memorization payoff, and no formal contamination study has been conducted, so contamination remains a consideration.
5 cited facts
This benchmark measures creative short-fiction quality through hard constraint satisfaction, requiring each story to weave in a set of required elements from auto-generated prompts of a target word count. Machine judges grading machine writing is the load-bearing assumption. Bias correction narrows the gap to human judgment without closing it, so read rankings as corrected judge preference rather than settled quality. The sharpest caveat is that, without any published human baseline, the scores reflect only relative LLM preferences and are vulnerable to self-preference and style-over-substance biases.
5 cited facts
Lech Mazur Writing (LLM Creative Story-Writing Benchmark): LLM Creative Story-Writing Benchmark: models write short stories that must organically incorporate 10 randomly-assigned required elements within a target word range, and other LLMs judge the outputs. The current canonical method (2026) is pairwise head-to-head comparison (relative Thurstone score); the widely-cited earlier method was a 0-10 LLM-grader rubric panel.
GPT-5 leads Lech Mazur Writing at 86. The full leaderboard above lists every recorded measurement, not just the headline number.
41 models have recorded scores on Lech Mazur Writing, spanning a score spread of 25.5.
tensor.news grades Lech Mazur Writing B for integrity (score 76/100), ranking #45 of 61 benchmarks we assess across discrimination, saturation, contamination resistance, harness comparability, and freshness.
No — Lech Mazur Writing still has headroom and continues to discriminate between models rather than bunching them at the ceiling.
Its test set is public, so contamination risk is on the table: a high score may partly reflect training-data overlap rather than capability.
Scores are reported under a consistent harness, so comparisons on Lech Mazur Writing are reasonably apples-to-apples.
Every Lech Mazur Writing measurement is source-backed and tagged reproduced, self-reported, or unverified; the underlying record cites github.com/lechmazur/writing.