Arena Multi-turn

In plain words · Which AI people preferred in longer back-and-forth conversations

What does it measure?
Preference when a conversation continues over several turns rather than a single exchange — whether the AI remembers earlier context and stays consistent. Close to the real chat experience.
Who evaluates it, and how?
On Arena Intelligence, a user enters a question, two anonymous AIs answer, and the user votes for the better one. Over 82 million votes are aggregated with a statistical model (Bradley-Terry) into an Elo-style score, and this table takes the 'style-controlled' version from the official dataset, which statistically removes the bias toward long, nicely formatted answers. Only conversations with two or more exchanges are counted.
Reading the score
This is a head-to-head record, not an exam score (accuracy %). It is computed from the probability of being preferred over other AIs, so there is no maximum, and when a new AI enters, existing scores shift slightly. Typical values are 1,000–1,500; a 100-point gap means the leader is preferred about 64% of the time, a 200-point gap about 76%. Correctness is not counted — a wrong but friendly answer can win — so if accuracy matters, also look at accuracy tests (GPQA, AIME, etc.). Gaps of a few dozen points are effectively the same tier even if the rank differs. Useful for separating AIs that only give a good first answer from those that stay coherent as the conversation grows.

Last updated: 2026-10-02

Rank

Top score1554Gemini 4 Argon
Models tested97
Last updated2026.10.02
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 🥇GoogleGemini 4 Argon15221554—
  2. 🥈MetaMuse Spark 1.214701521—
  3. 🥉AnthropicClaude Fable 51508151863.5%
  4. #4AnthropicClaude Opus 4.71484151458.6%
  5. #5AnthropicClaude Opus 4.61490151253.1%
  6. #6GoogleGemini 3.8 Flash14871501—
  7. #7AnthropicClaude Opus 5.515191500—
  8. #7Moonshot AIKimi K314761500—
  9. #9GoogleGemini 3.1 Pro1480149777.1%
  10. #10MetaMuse Spark 1.114611496—
  11. #11AnthropicClaude Opus 4.81468149462.2%
  12. #12MetaMuse Spark1463149375.9%
  13. #13OpenAIGPT-5.41465149373.9%
  14. #14GoogleGemini 3.7 Flash14961493—
  15. #15OpenAIGPT-6.1 Sol14721493—
  16. #16MetaMuse Spark 1.314701490—
  17. #17OpenAIGPT-6 Astra14651488—
  18. #18GoogleGemini 3.6 Flash14741488—
  19. #19AnthropicClaude Fable 5.114911488—
  20. #20OpenAIGPT-5.6 Sol1482148772.7%
  21. #21AnthropicClaude Opus 514801487—
  22. #22AnthropicClaude Opus 4.51467148458.0%
  23. #23OpenAIGPT-5.51466148475.9%
  24. #24xAIGrok 4.201452148349.3%
  25. #25AlibabaQwen3.7 Max1459148380.5%
  26. #26GoogleGemini 3 Flash1461148278.0%
  27. #27AnthropicClaude Sonnet 4.61455148156.6%
  28. #28XiaomiMiMo V2.5 Pro1450148179.9%
  29. #29Z.aiGLM 5.314661480—
  30. #30xAIGrok 4.20 (Reasoning)1448147981.2%
  31. #31Z.aiGLM-5.11452147876.3%
  32. #32GoogleGemini 3.5 Flash1472147876.3%
  33. #33AnthropicClaude Sonnet 4.51454147857.3%
  34. #34Z.aiGLM 5.21459147573.3%
  35. #34DeepSeekDeepSeek V4 Pro1451147576.5%
  36. #36AnthropicClaude Sonnet 5.514611473—
  37. #37AnthropicClaude Sonnet 514521473—
  38. #38xAIGrok 4.514581473—
  39. #39AnthropicClaude Opus 4.11446147255.4%
  40. #40Z.aiGLM-51449147172.3%
  41. #40Z.aiGLM 5.3 Flash14481471—
  42. #42OpenAIGPT-5.6 Terra1445147071.2%
  43. #43OpenAIGPT-6 Sol14391469—
  44. #44AlibabaQwen3.6 Max1446146976.6%
  45. #45XiaomiMiMo V2 Pro1431146668.8%
  46. #46OpenAIGPT-5.4 Mini1423146673.3%
  47. #47GoogleGemini 3.5 Flash-Lite14431466—
  48. #48GoogleGemma 4 31B1435146575.6%
  49. #49AlibabaQwen3.7 Plus1437146478.0%
  50. #50DeepSeekDeepSeek V4.1 Flash14541461—
  51. #51Moonshot AIKimi K2.61441146176.0%
  52. #52StepfunStep-514361460—
  53. #53XiaomiMiMo-V2.6-Flash14191457—
  54. #54OpenAIGPT-5.6 Luna14331457—
  55. #55XiaomiMiMo-V2.6-Pro14591456—
  56. #56AlibabaQwen3.5 397B A17B1421145478.8%
  57. #57DeepSeekDeepSeek V4 Flash1416145379.2%
  58. #58Moonshot AIKimi K2.51428145270.2%
  59. #59XiaomiMiMo V2.51412145167.1%
  60. #60MiniMaxMiniMax M31422145082.9%
  61. #61ByteDanceDola Seed 2.0 Pro14091449—
  62. #62GoogleGemini 2.5 Pro1445144848.7%
  63. #63OpenAIGPT-6 Luna14211448—
  64. #64xAIGrok 4.31424144683.3%
  65. #65AlibabaQwen3.6 Plus1423144575.2%
  66. #66xAIGrok 4.614461444—
  67. #67BaiduERNIE 5.0 Thinking1419144241.4%
  68. #68AnthropicClaude Opus 41431143953.7%
  69. #69GoogleGemini 3.1 Flash Lite1420143577.2%
  70. #70xAIGrok 4.714311433—
  71. #71Mistral AIMistral Medium 3.51405143368.8%
  72. #72Z.aiGLM 5V Turbo1414143261.1%
  73. #72MeituanLongcat Flash Chat1392143243.1%
  74. #74DeepSeekDeepSeek V3.21410142960.7%
  75. #75MiniMaxMiniMax M2.71387142675.7%
  76. #76AnthropicClaude Haiku 4.51400142554.3%
  77. #77AnthropicClaude Sonnet 41398142154.7%
  78. #78OpenAIGPT-51396142073.1%
  79. #79TencentHy31377141563.1%
  80. #80xAIGrok 4.1 Fast (Reasoning)1408141452.7%
  81. #81OpenAIGPT-5.4 Nano1363141375.9%
  82. #82GoogleGemini 2.5 Flash1402140350.3%
  83. #83NVIDIANemotron 3 Ultra1392139881.4%
  84. #84MiniMaxMiniMax M2.51369139771.6%
  85. #85UpstageSolar Pro 413401390—
  86. #86GoogleGemini 2.5 Flash Lite1371137452.6%
  87. #87OpenAIGPT-5 Mini1352137375.4%
  88. #88NVIDIANemotron 3 Super1325135171.5%
  89. #89Arcee AITrinity Large Thinking1343134656.3%
  90. #90OpenAIGPT OSS 120B1309132769.0%
  91. #91MetaLlama 4 Maverick1312132643.0%
  92. #92OpenAIGPT-5 Nano1289132467.6%
  93. #93AmazonNova 2 Lite1298132370.7%
  94. #94MetaLlama 4 Scout1305131939.5%
  95. #95OpenAIGPT-4.11308129943.0%
  96. #96NVIDIANemotron 3 Nano 30B A3B1267129071.1%
  97. #97IBMGranite 4.1 8B1274127938.6%
Half of models ≤ 1464Measured by: User preference votesLast updated 2026-10-02

18 of the 37 AIs released in the last 60 days have no score here yet · Claude Haiku 5.5, Decider V1.1 27B, GPT-6 Luna Decisions and 15 more

Source: Arena Intelligence Data sources & removal requests

What to keep in mind

  • It’s a preference vote, not a correctness check. A wrong but friendly answer can win.
  • Which questions got asked depends on who voted — they may differ from your use.
  • Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.