Arena WebDev
In plain words · Which AI built the websites and apps people preferred
- What does it measure?
- A ranking of which AI builds the more usable result when asked to make a website or app. People pick by looking at the running page rather than grading code, so a high score means the AI is good at producing something that looks right and works.
- Who evaluates it, and how?
- In the WebDev section of Arena Intelligence, a user describes a page or app, two anonymous AIs write the code, and the results are shown running side by side. The user votes for the better one, and votes are scored with the same statistical model as text Arena.
- Reading the score
- This is a head-to-head record, not an exam score. It is computed from the probability of being preferred over other AIs, so there is no maximum; a 100-point gap means the leader is preferred about 64% of the time. Gaps of a few dozen points are effectively the same tier even if the rank differs. Values in this table are currently about 1,200–1,800. First impressions of the running page drive the vote, so code maintainability and security are not reflected.
Last updated: 2026-10-01
Rank
Top score1788GPT-6 Astra
Models tested64
Last updated2026.10.01
Model release date (newest on the right)
Best score so farHigher on the chart is better
Can it operate a computer itself to finish hours-long development jobs?Terminal-Bench 4.0
Claude Sonnet 5.5🥇
Claude Sonnet 5.5🥇Which AI gives the coding answers people prefer most?Arena Coding
Gemini 4 Argon🥇
Gemini 4 Argon🥇Can it run and fix code until the task is done?LiveBench Agentic Coding
DeepSeek V4.1 Flash🥇
DeepSeek V4.1 Flash🥇
- 🥇OpenAIGPT-6 Astra1788
- 🥈AnthropicClaude Sonnet 5.51786
- 🥉OpenAIGPT-6.1 Sol1758
- #4AnthropicClaude Fable 5.11749
- #5AnthropicClaude Opus 51695
- #6OpenAIGPT-6 Sol1689
- #7GoogleGemini 4 Argon1680
- #8Moonshot AIKimi K31658
- #9MetaMuse Spark 1.31657
- #10xAIGrok 4.71638
- #11AnthropicClaude Fable 51625
- #12Z.aiGLM 5.31623
- #13DeepSeekDeepSeek V4.1 Flash1620
- #14xAIGrok 4.61620
- #15XiaomiMiMo-V2.6-Pro1618
- #16Z.aiGLM 5.3 Flash1616
- #17Z.aiGLM 5.21605
- #18GoogleGemini 3.7 Flash1592
- #19GoogleGemini 3.8 Flash1583
- #20OpenAIGPT-6 Luna1579
- #21AnthropicClaude Opus 4.71555
- #22xAIGrok 4.51552
- #23MetaMuse Spark 1.11542
- #24AnthropicClaude Sonnet 51539
- #25AnthropicClaude Opus 4.61537
- #26GoogleGemini 3.6 Flash1536
- #27AnthropicClaude Opus 4.81534
- #28MetaMuse Spark 1.21532
- #29AnthropicClaude Sonnet 4.61521
- #30Moonshot AIKimi K2.61509
- #31Z.aiGLM-5.11509
- #32GoogleGemini 3.5 Flash1490
- #33MiniMaxMiniMax M31482
- #34AlibabaQwen3.6 Max1482
- #35XiaomiMiMo V2.5 Pro1478
- #36AnthropicClaude Opus 4.51469
- #37AlibabaQwen3.6 Plus1461
- #38GoogleGemini 3.1 Pro1446
- #39DeepSeekDeepSeek V4 Pro1446
- #40GoogleGemini 3.5 Flash-Lite1440
- #41GoogleGemini 3 Flash1439
- #42XiaomiMiMo V2.51438
- #43Moonshot AIKimi K2.51436
- #44Z.aiGLM-51434
- #45XiaomiMiMo V2 Pro1433
- #46Z.aiGLM 5V Turbo1400
- #47AlibabaQwen3.5 397B A17B1399
- #48OpenAIGPT-5.4 Mini1397
- #48MiniMaxMiniMax M2.71397
- #50MiniMaxMiniMax M2.51386
- #51AnthropicClaude Sonnet 4.51384
- #52xAIGrok 4.20 (Reasoning)1374
- #53UpstageSolar Pro 41370
- #54GoogleGemma 4 31B1365
- #55xAIGrok 4.31357
- #56TencentHy31356
- #57AnthropicClaude Haiku 4.51329
- #58DeepSeekDeepSeek V3.21325
- #59Mistral AIMistral Medium 3.51263
- #60GoogleGemini 3.1 Flash Lite1256
- #61xAIGrok 4.1 Fast (Reasoning)1241
- #62Arcee AITrinity Large Thinking1238
- #63GoogleGemini 2.5 Pro1227
- #64IBMGranite 4.1 8B1190
Half of models ≤ 1486Measured by: User preference votesLast updated 2026-10-01
18 of the 34 AIs released in the last 60 days have no score here yet · Mistral Large 4, Clef, Perplexity Decider v1 27B and 15 more
Source: Arena Intelligence Data sources & removal requestsWhat to keep in mind
- It’s a preference vote, not a check that the request was followed. A result that looks or sounds nicer can win even if it strays.
- Which requests got made depends on who voted — they may differ from your use.
- Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.