Arena Text

In plain words · Which AI people preferred across every kind of question

What does it measure?
A popularity ranking of which AI's answers people preferred across every kind of question they actually asked. It is not an exam score but 'user preference', so it is the closest thing to real-world satisfaction.
Who evaluates it, and how?
On Arena Intelligence, a user enters a question, two anonymous AIs answer, and the user votes for the better one. Over 82 million votes are aggregated with a statistical model (Bradley-Terry) into an Elo-style score, and this table takes the 'style-controlled' version from the official dataset, which statistically removes the bias toward long, nicely formatted answers. This table is the combined score across all kinds of questions.
Reading the score
This is a head-to-head record, not an exam score (accuracy %). It is computed from the probability of being preferred over other AIs, so there is no maximum, and when a new AI enters, existing scores shift slightly. Typical values are 1,000–1,500; a 100-point gap means the leader is preferred about 64% of the time, a 200-point gap about 76%. Correctness is not counted — a wrong but friendly answer can win — so if accuracy matters, also look at accuracy tests (GPQA, AIME, etc.). Gaps of a few dozen points are effectively the same tier even if the rank differs.

Last updated: 2026-10-02

Try it yourself

Pick AIs and compare them.

Some AIs cannot be selected.

Rank

Top score1525Gemini 4 Argon
Models tested97
Last updated2026.10.02
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 🥇GoogleGemini 4 Argon52.61525—
  2. 🥈AnthropicClaude Fable 549.6150483.0
  3. 🥉AnthropicClaude Opus 5.557.6150482.1
  4. #4AnthropicClaude Fable 5.153.4150183.4
  5. #5AnthropicClaude Opus 4.631.9149774.5
  6. #6GoogleGemini 3.8 Flash40.9149575.8
  7. #7MetaMuse Spark 1.348.1149481.6
  8. #8AnthropicClaude Opus 4.740.7149476.5
  9. #9MetaMuse Spark 1.239.6149478.0
  10. #10MetaMuse Spark 1.133.7149175.3
  11. #11AnthropicClaude Opus 550.8149080.1
  12. #12MetaMuse Spark31.31489—
  13. #13Moonshot AIKimi K343.6148879.2
  14. #14GoogleGemini 3.7 Flash39.6148878.8
  15. #15GoogleGemini 3.1 Pro29.7148777.0
  16. #16OpenAIGPT-5.6 Sol47.0148481.0
  17. #17OpenAIGPT-6.1 Sol51.8148381.1
  18. #18GoogleGemini 3.6 Flash34.0148373.6
  19. #19XiaomiMiMo-V2.6-Pro46.31480—
  20. #20Z.aiGLM 5.344.8147976.1
  21. #21OpenAIGPT-6 Astra52.7147782.2
  22. #22OpenAIGPT-5.538.4147780.2
  23. #23GoogleGemini 3.5 Flash33.6147674.6
  24. #24Z.aiGLM 5.233.7147673.2
  25. #25OpenAIGPT-5.439.0147578.0
  26. #26AlibabaQwen3.7 Max29.5147573.1
  27. #27xAIGrok 4.2014.21475—
  28. #28AnthropicClaude Opus 4.841.8147576.2
  29. #29DeepSeekDeepSeek V4.1 Flash39.5147481.1
  30. #30Z.aiGLM 5.3 Flash41.8147371.6
  31. #31GoogleGemini 3 Flash26.31473—
  32. #32AnthropicClaude Sonnet 4.630.1147273.0
  33. #33xAIGrok 4.20 (Reasoning)25.71472—
  34. #34AnthropicClaude Sonnet 5.556.0147177.8
  35. #35AnthropicClaude Opus 4.529.1147072.6
  36. #36XiaomiMiMo V2.5 Pro26.01468—
  37. #37xAIGrok 4.538.8146675.8
  38. #38OpenAIGPT-5.6 Terra42.1146677.9
  39. #39Z.aiGLM-5.126.11465—
  40. #40AnthropicClaude Sonnet 538.2146276.0
  41. #41Moonshot AIKimi K2.627.0146170.5
  42. #42AlibabaQwen3.6 Max28.41460—
  43. #43DeepSeekDeepSeek V4 Pro36.0145871.6
  44. #44Z.aiGLM-527.91458—
  45. #45OpenAIGPT-6 Sol47.6145779.3
  46. #46ByteDanceDola Seed 2.0 Pro—1457—
  47. #47AlibabaQwen3.7 Plus25.21456—
  48. #48AnthropicClaude Sonnet 4.520.71455—
  49. #49GoogleGemini 3.5 Flash-Lite22.2145563.9
  50. #50xAIGrok 4.644.3145478.0
  51. #51StepfunStep-543.71454—
  52. #52OpenAIGPT-5.6 Luna37.3145373.6
  53. #53GoogleGemma 4 31B14.71453—
  54. #54XiaomiMiMo-V2.6-Flash37.91452—
  55. #55Moonshot AIKimi K2.523.51450—
  56. #56AnthropicClaude Opus 4.122.81450—
  57. #57XiaomiMiMo V2 Pro28.61448—
  58. #58OpenAIGPT-5.4 Mini24.1144766.4
  59. #59BaiduERNIE 5.0 Thinking14.31446—
  60. #60GoogleGemini 2.5 Pro16.11446—
  61. #61AlibabaQwen3.6 Plus27.0144368.9
  62. #62OpenAIGPT-6 Luna38.1144372.0
  63. #63xAIGrok 4.746.4144377.4
  64. #64AlibabaQwen3.5 397B A17B21.41442—
  65. #65xAIGrok 4.324.9144162.3
  66. #66MiniMaxMiniMax M329.2144067.3
  67. #67MeituanLongcat Flash Chat11.51437—
  68. #68DeepSeekDeepSeek V4 Flash34.3143665.5
  69. #69OpenAIGPT-523.01435—
  70. #70XiaomiMiMo V2.525.21434—
  71. #71Z.aiGLM 5V Turbo23.51434—
  72. #72GoogleGemini 3.1 Flash Lite15.61433—
  73. #73xAIGrok 4.1 Fast (Reasoning)20.41430—
  74. #74Mistral AIMistral Medium 3.514.21427—
  75. #75NVIDIANemotron 3 Ultra22.9142667.4
  76. #76AnthropicClaude Opus 420.61426—
  77. #77DeepSeekDeepSeek V3.221.51425—
  78. #78MiniMaxMiniMax M2.722.81415—
  79. #79AnthropicClaude Haiku 4.516.91414—
  80. #80GoogleGemini 2.5 Flash13.11409—
  81. #81TencentHy325.31409—
  82. #82AnthropicClaude Sonnet 418.91402—
  83. #83OpenAIGPT-5.4 Nano20.7140169.6
  84. #84MiniMaxMiniMax M2.522.81391—
  85. #85OpenAIGPT-5 Mini20.61390—
  86. #86UpstageSolar Pro 428.21386—
  87. #87GoogleGemini 2.5 Flash Lite10.41379—
  88. #88Arcee AITrinity Large Thinking10.81367—
  89. #89NVIDIANemotron 3 Super12.81361—
  90. #90OpenAIGPT OSS 120B11.61352—
  91. #91OpenAIGPT-5 Nano13.01338—
  92. #92AmazonNova 2 Lite13.41335—
  93. #93MetaLlama 4 Maverick10.01327—
  94. #94MetaLlama 4 Scout8.11322—
  95. #95NVIDIANemotron 3 Nano 30B A3B8.91314—
  96. #96OpenAIGPT-4.112.71313—
  97. #97IBMGranite 4.1 8B6.61304—
Half of models ≤ 1455Measured by: User preference votesLast updated 2026-10-02

18 of the 37 AIs released in the last 60 days have no score here yet · Claude Haiku 5.5, Decider V1.1 27B, GPT-6 Luna Decisions and 15 more

Source: Arena Intelligence Data sources & removal requests

What to keep in mind

  • It’s a preference vote, not a correctness check. A wrong but friendly answer can win.
  • Which questions got asked depends on who voted — they may differ from your use.
  • A hands-on comparison here is two AIs and one question, so it can differ from the official ranking. More rounds make it more accurate.
  • Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.