Arena Creative Writing
In plain words · Which AI people preferred for creative writing
- What does it measure?
- Preference on writing that needs imagination and expression — fiction, poetry, essays, columns. It reveals how people judged style, voice and emotional expression.
- Who evaluates it, and how?
- On Arena Intelligence, a user enters a question, two anonymous AIs answer, and the user votes for the better one. Over 82 million votes are aggregated with a statistical model (Bradley-Terry) into an Elo-style score, and this table takes the 'style-controlled' version from the official dataset, which statistically removes the bias toward long, nicely formatted answers. An AI classifier trained on human-labelled examples selects prompts asking for originality, emotional expression or a point of view.
- Reading the score
- This is a head-to-head record, not an exam score (accuracy %). It is computed from the probability of being preferred over other AIs, so there is no maximum, and when a new AI enters, existing scores shift slightly. Typical values are 1,000–1,500; a 100-point gap means the leader is preferred about 64% of the time, a 200-point gap about 76%. Correctness is not counted — a wrong but friendly answer can win — so if accuracy matters, also look at accuracy tests (GPQA, AIME, etc.). Gaps of a few dozen points are effectively the same tier even if the rank differs. Tastes vary here, so voter preferences influence this category more than others.
Last updated: 2026-10-02
Rank
Top score1519Gemini 4 Argon
Models tested97
Last updated2026.10.02
Model release date (newest on the right)
Best score so farHigher on the chart is better
Which AI gives the writing and literature answers people prefer most?Arena Writing, Literature & Language
Gemini 4 Argon🥇
Gemini 4 Argon🥇Which AI do people prefer most in long conversations?Arena Multi-turn
Gemini 4 Argon🥇
Gemini 4 Argon🥇Does it follow tricky rules like format and length without missing any?IFBench
Grok 4.3🥇
Grok 4.3🥇
- 🥇GoogleGemini 4 Argon1519
- 🥈AnthropicClaude Opus 5.51516
- 🥉AnthropicClaude Fable 51503
- #4GoogleGemini 3.7 Flash1493
- #5GoogleGemini 3.8 Flash1485
- #6AnthropicClaude Opus 4.71483
- #7AnthropicClaude Fable 5.11482
- #8GoogleGemini 3.1 Pro1480
- #9AnthropicClaude Opus 4.61479
- #10GoogleGemini 3.6 Flash1471
- #11AnthropicClaude Opus 51471
- #12GoogleGemini 3.5 Flash1469
- #13OpenAIGPT-5.6 Sol1468
- #14MetaMuse Spark1466
- #15AnthropicClaude Opus 4.81463
- #15xAIGrok 4.201463
- #17AnthropicClaude Opus 4.51462
- #18Moonshot AIKimi K31460
- #19OpenAIGPT-6.1 Sol1460
- #20MetaMuse Spark 1.31459
- #21GoogleGemini 3 Flash1457
- #22MetaMuse Spark 1.21456
- #23Z.aiGLM 5.21454
- #24AnthropicClaude Sonnet 4.51454
- #25Z.aiGLM 5.31454
- #26xAIGrok 4.51450
- #27AnthropicClaude Sonnet 4.61449
- #28OpenAIGPT-6 Astra1449
- #29Z.aiGLM-51449
- #30OpenAIGPT-5.51448
- #31AnthropicClaude Sonnet 5.51447
- #32MetaMuse Spark 1.11447
- #33Z.aiGLM-5.11446
- #34XiaomiMiMo-V2.6-Pro1446
- #35DeepSeekDeepSeek V4 Pro1446
- #36AnthropicClaude Opus 4.11446
- #37xAIGrok 4.20 (Reasoning)1446
- #38xAIGrok 4.61445
- #39GoogleGemini 2.5 Pro1443
- #40OpenAIGPT-5.41443
- #41OpenAIGPT-6 Sol1442
- #42AlibabaQwen3.7 Max1441
- #43DeepSeekDeepSeek V4.1 Flash1440
- #44AlibabaQwen3.6 Max1437
- #45GoogleGemini 3.5 Flash-Lite1436
- #46xAIGrok 4.71436
- #47AnthropicClaude Sonnet 51435
- #48XiaomiMiMo V2.5 Pro1434
- #49Z.aiGLM 5.3 Flash1433
- #50Moonshot AIKimi K2.61433
- #51AnthropicClaude Opus 41432
- #52AlibabaQwen3.7 Plus1432
- #53OpenAIGPT-5.6 Terra1426
- #54XiaomiMiMo V2 Pro1424
- #55Moonshot AIKimi K2.51423
- #56xAIGrok 4.31422
- #57BaiduERNIE 5.0 Thinking1422
- #58StepfunStep-51421
- #59GoogleGemma 4 31B1420
- #60xAIGrok 4.1 Fast (Reasoning)1414
- #61GoogleGemini 3.1 Flash Lite1413
- #62OpenAIGPT-6 Luna1413
- #63OpenAIGPT-5.6 Luna1409
- #64DeepSeekDeepSeek V4 Flash1408
- #65AlibabaQwen3.6 Plus1407
- #66MiniMaxMiniMax M31406
- #67Z.aiGLM 5V Turbo1406
- #68AlibabaQwen3.5 397B A17B1405
- #69ByteDanceDola Seed 2.0 Pro1403
- #70OpenAIGPT-5.4 Mini1401
- #71DeepSeekDeepSeek V3.21401
- #72Mistral AIMistral Medium 3.51400
- #73AnthropicClaude Sonnet 41396
- #74XiaomiMiMo V2.51394
- #75GoogleGemini 2.5 Flash1394
- #76MeituanLongcat Flash Chat1390
- #77AnthropicClaude Haiku 4.51388
- #78XiaomiMiMo-V2.6-Flash1385
- #79NVIDIANemotron 3 Ultra1382
- #80OpenAIGPT-51373
- #81MiniMaxMiniMax M2.71364
- #82GoogleGemini 2.5 Flash Lite1360
- #83MiniMaxMiniMax M2.51357
- #84TencentHy31353
- #85OpenAIGPT-5.4 Nano1336
- #86Arcee AITrinity Large Thinking1332
- #87OpenAIGPT-5 Mini1324
- #88UpstageSolar Pro 41310
- #89MetaLlama 4 Maverick1307
- #90NVIDIANemotron 3 Super1303
- #91MetaLlama 4 Scout1289
- #92OpenAIGPT-4.11286
- #93OpenAIGPT OSS 120B1278
- #94AmazonNova 2 Lite1269
- #95IBMGranite 4.1 8B1262
- #96OpenAIGPT-5 Nano1247
- #97NVIDIANemotron 3 Nano 30B A3B1242
Half of models ≤ 1433Measured by: User preference votesLast updated 2026-10-02
15 of the 34 AIs released in the last 60 days have no score here yet · Mistral Large 4, Perplexity Decider v1 27B, Clef Flash and 12 more
Source: Arena Intelligence Data sources & removal requestsWhat to keep in mind
- It’s a preference vote, not a correctness check. A wrong but friendly answer can win.
- Which questions got asked depends on who voted — they may differ from your use.
- Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.