Lech Mazur Creative Writing
Higher is better
Lech Mazur's creative-writing test: models write short stories that must naturally include ten randomly assigned elements (character, object, setting, motive, and so on), rated 0 to 10 by a panel of seven AI judges. Values here are frozen at the old scoring method as of August 2025. Scores are shown as published by the source. Higher is better.
Top score8.60 / 10GPT-5
Models tested8
Model release date (newest on the right)
Best score so farHigher on the chart is better
Which AI gives the writing and literature answers people prefer most?Arena Writing, Literature & Language
Gemini 4 Argon🥇
Gemini 4 Argon🥇Which AI do people prefer most in long conversations?Arena Multi-turn
Gemini 4 Argon🥇
Gemini 4 Argon🥇Does it follow tricky rules like format and length without missing any?IFBench
Grok 4.3🥇
Grok 4.3🥇
- 🥇OpenAIGPT-5Reasoning effort: Medium8.60 / 10
- 🥈AnthropicClaude Opus 4.18.47 / 10
- 🥉GoogleGemini 2.5 Pro8.38 / 10
- #4AnthropicClaude Opus 4Reasoning effort: Budget 16K8.36 / 10
- #5OpenAIGPT-5 MiniReasoning effort: Medium8.31 / 10
- #6AnthropicClaude Sonnet 4Reasoning effort: Budget 16K8.14 / 10
- #7OpenAIGPT OSS 120B7.71 / 10
- #8MetaLlama 4 Maverick6.20 / 10
Half of models ≤ 8.34 / 10Last updated 2026-10-07
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.