Arena Multi-turn
In plain words · Which AI people preferred in longer back-and-forth conversations
- What does it measure?
- Preference when a conversation continues over several turns rather than a single exchange — whether the AI remembers earlier context and stays consistent. Close to the real chat experience.
- Who evaluates it, and how?
- On Arena Intelligence, a user enters a question, two anonymous AIs answer, and the user votes for the better one. Over 82 million votes are aggregated with a statistical model (Bradley-Terry) into an Elo-style score, and this table takes the 'style-controlled' version from the official dataset, which statistically removes the bias toward long, nicely formatted answers. Only conversations with two or more exchanges are counted.
- Reading the score
- This is a head-to-head record, not an exam score (accuracy %). It is computed from the probability of being preferred over other AIs, so there is no maximum, and when a new AI enters, existing scores shift slightly. Typical values are 1,000–1,500; a 100-point gap means the leader is preferred about 64% of the time, a 200-point gap about 76%. Correctness is not counted — a wrong but friendly answer can win — so if accuracy matters, also look at accuracy tests (GPQA, AIME, etc.). Gaps of a few dozen points are effectively the same tier even if the rank differs. Useful for separating AIs that only give a good first answer from those that stay coherent as the conversation grows.
Last updated: 2026-10-02
Rank
Top score1554Gemini 4 Argon
Models tested97
Last updated2026.10.02
Model release date (newest on the right)
Best score so farHigher on the chart is better
Which AI gives the writing and literature answers people prefer most?Arena Writing, Literature & Language
Gemini 4 Argon🥇
Gemini 4 Argon🥇Does it follow tricky rules like format and length without missing any?IFBench
Grok 4.3🥇
Grok 4.3🥇
- 🥇GoogleGemini 4 Argon15221554—
- 🥈MetaMuse Spark 1.214701521—
- 🥉AnthropicClaude Fable 51508151863.5%
- #4AnthropicClaude Opus 4.71484151458.6%
- #5AnthropicClaude Opus 4.61490151253.1%
- #6GoogleGemini 3.8 Flash14871501—
- #7AnthropicClaude Opus 5.515191500—
- #7Moonshot AIKimi K314761500—
- #9GoogleGemini 3.1 Pro1480149777.1%
- #10MetaMuse Spark 1.114611496—
- #11AnthropicClaude Opus 4.81468149462.2%
- #12MetaMuse Spark1463149375.9%
- #13OpenAIGPT-5.41465149373.9%
- #14GoogleGemini 3.7 Flash14961493—
- #15OpenAIGPT-6.1 Sol14721493—
- #16MetaMuse Spark 1.314701490—
- #17OpenAIGPT-6 Astra14651488—
- #18GoogleGemini 3.6 Flash14741488—
- #19AnthropicClaude Fable 5.114911488—
- #20OpenAIGPT-5.6 Sol1482148772.7%
- #21AnthropicClaude Opus 514801487—
- #22AnthropicClaude Opus 4.51467148458.0%
- #23OpenAIGPT-5.51466148475.9%
- #24xAIGrok 4.201452148349.3%
- #25AlibabaQwen3.7 Max1459148380.5%
- #26GoogleGemini 3 Flash1461148278.0%
- #27AnthropicClaude Sonnet 4.61455148156.6%
- #28XiaomiMiMo V2.5 Pro1450148179.9%
- #29Z.aiGLM 5.314661480—
- #30xAIGrok 4.20 (Reasoning)1448147981.2%
- #31Z.aiGLM-5.11452147876.3%
- #32GoogleGemini 3.5 Flash1472147876.3%
- #33AnthropicClaude Sonnet 4.51454147857.3%
- #34Z.aiGLM 5.21459147573.3%
- #34DeepSeekDeepSeek V4 Pro1451147576.5%
- #36AnthropicClaude Sonnet 5.514611473—
- #37AnthropicClaude Sonnet 514521473—
- #38xAIGrok 4.514581473—
- #39AnthropicClaude Opus 4.11446147255.4%
- #40Z.aiGLM-51449147172.3%
- #40Z.aiGLM 5.3 Flash14481471—
- #42OpenAIGPT-5.6 Terra1445147071.2%
- #43OpenAIGPT-6 Sol14391469—
- #44AlibabaQwen3.6 Max1446146976.6%
- #45XiaomiMiMo V2 Pro1431146668.8%
- #46OpenAIGPT-5.4 Mini1423146673.3%
- #47GoogleGemini 3.5 Flash-Lite14431466—
- #48GoogleGemma 4 31B1435146575.6%
- #49AlibabaQwen3.7 Plus1437146478.0%
- #50DeepSeekDeepSeek V4.1 Flash14541461—
- #51Moonshot AIKimi K2.61441146176.0%
- #52StepfunStep-514361460—
- #53XiaomiMiMo-V2.6-Flash14191457—
- #54OpenAIGPT-5.6 Luna14331457—
- #55XiaomiMiMo-V2.6-Pro14591456—
- #56AlibabaQwen3.5 397B A17B1421145478.8%
- #57DeepSeekDeepSeek V4 Flash1416145379.2%
- #58Moonshot AIKimi K2.51428145270.2%
- #59XiaomiMiMo V2.51412145167.1%
- #60MiniMaxMiniMax M31422145082.9%
- #61ByteDanceDola Seed 2.0 Pro14091449—
- #62GoogleGemini 2.5 Pro1445144848.7%
- #63OpenAIGPT-6 Luna14211448—
- #64xAIGrok 4.31424144683.3%
- #65AlibabaQwen3.6 Plus1423144575.2%
- #66xAIGrok 4.614461444—
- #67BaiduERNIE 5.0 Thinking1419144241.4%
- #68AnthropicClaude Opus 41431143953.7%
- #69GoogleGemini 3.1 Flash Lite1420143577.2%
- #70xAIGrok 4.714311433—
- #71Mistral AIMistral Medium 3.51405143368.8%
- #72Z.aiGLM 5V Turbo1414143261.1%
- #72MeituanLongcat Flash Chat1392143243.1%
- #74DeepSeekDeepSeek V3.21410142960.7%
- #75MiniMaxMiniMax M2.71387142675.7%
- #76AnthropicClaude Haiku 4.51400142554.3%
- #77AnthropicClaude Sonnet 41398142154.7%
- #78OpenAIGPT-51396142073.1%
- #79TencentHy31377141563.1%
- #80xAIGrok 4.1 Fast (Reasoning)1408141452.7%
- #81OpenAIGPT-5.4 Nano1363141375.9%
- #82GoogleGemini 2.5 Flash1402140350.3%
- #83NVIDIANemotron 3 Ultra1392139881.4%
- #84MiniMaxMiniMax M2.51369139771.6%
- #85UpstageSolar Pro 413401390—
- #86GoogleGemini 2.5 Flash Lite1371137452.6%
- #87OpenAIGPT-5 Mini1352137375.4%
- #88NVIDIANemotron 3 Super1325135171.5%
- #89Arcee AITrinity Large Thinking1343134656.3%
- #90OpenAIGPT OSS 120B1309132769.0%
- #91MetaLlama 4 Maverick1312132643.0%
- #92OpenAIGPT-5 Nano1289132467.6%
- #93AmazonNova 2 Lite1298132370.7%
- #94MetaLlama 4 Scout1305131939.5%
- #95OpenAIGPT-4.11308129943.0%
- #96NVIDIANemotron 3 Nano 30B A3B1267129071.1%
- #97IBMGranite 4.1 8B1274127938.6%
Half of models ≤ 1464Measured by: User preference votesLast updated 2026-10-02
18 of the 37 AIs released in the last 60 days have no score here yet · Claude Haiku 5.5, Decider V1.1 27B, GPT-6 Luna Decisions and 15 more
Source: Arena Intelligence Data sources & removal requestsWhat to keep in mind
- It’s a preference vote, not a correctness check. A wrong but friendly answer can win.
- Which questions got asked depends on who voted — they may differ from your use.
- Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.