Arena Creative Writing

In plain words · Which AI people preferred for creative writing

What does it measure?
Preference on writing that needs imagination and expression — fiction, poetry, essays, columns. It reveals how people judged style, voice and emotional expression.
Who evaluates it, and how?
On Arena Intelligence, a user enters a question, two anonymous AIs answer, and the user votes for the better one. Over 82 million votes are aggregated with a statistical model (Bradley-Terry) into an Elo-style score, and this table takes the 'style-controlled' version from the official dataset, which statistically removes the bias toward long, nicely formatted answers. An AI classifier trained on human-labelled examples selects prompts asking for originality, emotional expression or a point of view.
Reading the score
This is a head-to-head record, not an exam score (accuracy %). It is computed from the probability of being preferred over other AIs, so there is no maximum, and when a new AI enters, existing scores shift slightly. Typical values are 1,000–1,500; a 100-point gap means the leader is preferred about 64% of the time, a 200-point gap about 76%. Correctness is not counted — a wrong but friendly answer can win — so if accuracy matters, also look at accuracy tests (GPQA, AIME, etc.). Gaps of a few dozen points are effectively the same tier even if the rank differs. Tastes vary here, so voter preferences influence this category more than others.

Last updated: 2026-10-02

Rank

Top score1519Gemini 4 Argon
Models tested97
Last updated2026.10.02
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 🥇GoogleGemini 4 Argon1519
  2. 🥈AnthropicClaude Opus 5.51516
  3. 🥉AnthropicClaude Fable 51503
  4. #4GoogleGemini 3.7 Flash1493
  5. #5GoogleGemini 3.8 Flash1485
  6. #6AnthropicClaude Opus 4.71483
  7. #7AnthropicClaude Fable 5.11482
  8. #8GoogleGemini 3.1 Pro1480
  9. #9AnthropicClaude Opus 4.61479
  10. #10GoogleGemini 3.6 Flash1471
  11. #11AnthropicClaude Opus 51471
  12. #12GoogleGemini 3.5 Flash1469
  13. #13OpenAIGPT-5.6 Sol1468
  14. #14MetaMuse Spark1466
  15. #15AnthropicClaude Opus 4.81463
  16. #15xAIGrok 4.201463
  17. #17AnthropicClaude Opus 4.51462
  18. #18Moonshot AIKimi K31460
  19. #19OpenAIGPT-6.1 Sol1460
  20. #20MetaMuse Spark 1.31459
  21. #21GoogleGemini 3 Flash1457
  22. #22MetaMuse Spark 1.21456
  23. #23Z.aiGLM 5.21454
  24. #24AnthropicClaude Sonnet 4.51454
  25. #25Z.aiGLM 5.31454
  26. #26xAIGrok 4.51450
  27. #27AnthropicClaude Sonnet 4.61449
  28. #28OpenAIGPT-6 Astra1449
  29. #29Z.aiGLM-51449
  30. #30OpenAIGPT-5.51448
  31. #31AnthropicClaude Sonnet 5.51447
  32. #32MetaMuse Spark 1.11447
  33. #33Z.aiGLM-5.11446
  34. #34XiaomiMiMo-V2.6-Pro1446
  35. #35DeepSeekDeepSeek V4 Pro1446
  36. #36AnthropicClaude Opus 4.11446
  37. #37xAIGrok 4.20 (Reasoning)1446
  38. #38xAIGrok 4.61445
  39. #39GoogleGemini 2.5 Pro1443
  40. #40OpenAIGPT-5.41443
  41. #41OpenAIGPT-6 Sol1442
  42. #42AlibabaQwen3.7 Max1441
  43. #43DeepSeekDeepSeek V4.1 Flash1440
  44. #44AlibabaQwen3.6 Max1437
  45. #45GoogleGemini 3.5 Flash-Lite1436
  46. #46xAIGrok 4.71436
  47. #47AnthropicClaude Sonnet 51435
  48. #48XiaomiMiMo V2.5 Pro1434
  49. #49Z.aiGLM 5.3 Flash1433
  50. #50Moonshot AIKimi K2.61433
  51. #51AnthropicClaude Opus 41432
  52. #52AlibabaQwen3.7 Plus1432
  53. #53OpenAIGPT-5.6 Terra1426
  54. #54XiaomiMiMo V2 Pro1424
  55. #55Moonshot AIKimi K2.51423
  56. #56xAIGrok 4.31422
  57. #57BaiduERNIE 5.0 Thinking1422
  58. #58StepfunStep-51421
  59. #59GoogleGemma 4 31B1420
  60. #60xAIGrok 4.1 Fast (Reasoning)1414
  61. #61GoogleGemini 3.1 Flash Lite1413
  62. #62OpenAIGPT-6 Luna1413
  63. #63OpenAIGPT-5.6 Luna1409
  64. #64DeepSeekDeepSeek V4 Flash1408
  65. #65AlibabaQwen3.6 Plus1407
  66. #66MiniMaxMiniMax M31406
  67. #67Z.aiGLM 5V Turbo1406
  68. #68AlibabaQwen3.5 397B A17B1405
  69. #69ByteDanceDola Seed 2.0 Pro1403
  70. #70OpenAIGPT-5.4 Mini1401
  71. #71DeepSeekDeepSeek V3.21401
  72. #72Mistral AIMistral Medium 3.51400
  73. #73AnthropicClaude Sonnet 41396
  74. #74XiaomiMiMo V2.51394
  75. #75GoogleGemini 2.5 Flash1394
  76. #76MeituanLongcat Flash Chat1390
  77. #77AnthropicClaude Haiku 4.51388
  78. #78XiaomiMiMo-V2.6-Flash1385
  79. #79NVIDIANemotron 3 Ultra1382
  80. #80OpenAIGPT-51373
  81. #81MiniMaxMiniMax M2.71364
  82. #82GoogleGemini 2.5 Flash Lite1360
  83. #83MiniMaxMiniMax M2.51357
  84. #84TencentHy31353
  85. #85OpenAIGPT-5.4 Nano1336
  86. #86Arcee AITrinity Large Thinking1332
  87. #87OpenAIGPT-5 Mini1324
  88. #88UpstageSolar Pro 41310
  89. #89MetaLlama 4 Maverick1307
  90. #90NVIDIANemotron 3 Super1303
  91. #91MetaLlama 4 Scout1289
  92. #92OpenAIGPT-4.11286
  93. #93OpenAIGPT OSS 120B1278
  94. #94AmazonNova 2 Lite1269
  95. #95IBMGranite 4.1 8B1262
  96. #96OpenAIGPT-5 Nano1247
  97. #97NVIDIANemotron 3 Nano 30B A3B1242
Half of models ≤ 1433Measured by: User preference votesLast updated 2026-10-02

15 of the 34 AIs released in the last 60 days have no score here yet · Mistral Large 4, Perplexity Decider v1 27B, Clef Flash and 12 more

Source: Arena Intelligence Data sources & removal requests

What to keep in mind

  • It’s a preference vote, not a correctness check. A wrong but friendly answer can win.
  • Which questions got asked depends on who voted — they may differ from your use.
  • Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.