Arena WebDev

In plain words · Which AI built the websites and apps people preferred

What does it measure?
A ranking of which AI builds the more usable result when asked to make a website or app. People pick by looking at the running page rather than grading code, so a high score means the AI is good at producing something that looks right and works.
Who evaluates it, and how?
In the WebDev section of Arena Intelligence, a user describes a page or app, two anonymous AIs write the code, and the results are shown running side by side. The user votes for the better one, and votes are scored with the same statistical model as text Arena.
Reading the score
This is a head-to-head record, not an exam score. It is computed from the probability of being preferred over other AIs, so there is no maximum; a 100-point gap means the leader is preferred about 64% of the time. Gaps of a few dozen points are effectively the same tier even if the rank differs. Values in this table are currently about 1,200–1,800. First impressions of the running page drive the vote, so code maintainability and security are not reflected.

Last updated: 2026-10-01

Rank

Top score1788GPT-6 Astra
Models tested64
Last updated2026.10.01
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 🥇OpenAIGPT-6 Astra1788
  2. 🥈AnthropicClaude Sonnet 5.51786
  3. 🥉OpenAIGPT-6.1 Sol1758
  4. #4AnthropicClaude Fable 5.11749
  5. #5AnthropicClaude Opus 51695
  6. #6OpenAIGPT-6 Sol1689
  7. #7GoogleGemini 4 Argon1680
  8. #8Moonshot AIKimi K31658
  9. #9MetaMuse Spark 1.31657
  10. #10xAIGrok 4.71638
  11. #11AnthropicClaude Fable 51625
  12. #12Z.aiGLM 5.31623
  13. #13DeepSeekDeepSeek V4.1 Flash1620
  14. #14xAIGrok 4.61620
  15. #15XiaomiMiMo-V2.6-Pro1618
  16. #16Z.aiGLM 5.3 Flash1616
  17. #17Z.aiGLM 5.21605
  18. #18GoogleGemini 3.7 Flash1592
  19. #19GoogleGemini 3.8 Flash1583
  20. #20OpenAIGPT-6 Luna1579
  21. #21AnthropicClaude Opus 4.71555
  22. #22xAIGrok 4.51552
  23. #23MetaMuse Spark 1.11542
  24. #24AnthropicClaude Sonnet 51539
  25. #25AnthropicClaude Opus 4.61537
  26. #26GoogleGemini 3.6 Flash1536
  27. #27AnthropicClaude Opus 4.81534
  28. #28MetaMuse Spark 1.21532
  29. #29AnthropicClaude Sonnet 4.61521
  30. #30Moonshot AIKimi K2.61509
  31. #31Z.aiGLM-5.11509
  32. #32GoogleGemini 3.5 Flash1490
  33. #33MiniMaxMiniMax M31482
  34. #34AlibabaQwen3.6 Max1482
  35. #35XiaomiMiMo V2.5 Pro1478
  36. #36AnthropicClaude Opus 4.51469
  37. #37AlibabaQwen3.6 Plus1461
  38. #38GoogleGemini 3.1 Pro1446
  39. #39DeepSeekDeepSeek V4 Pro1446
  40. #40GoogleGemini 3.5 Flash-Lite1440
  41. #41GoogleGemini 3 Flash1439
  42. #42XiaomiMiMo V2.51438
  43. #43Moonshot AIKimi K2.51436
  44. #44Z.aiGLM-51434
  45. #45XiaomiMiMo V2 Pro1433
  46. #46Z.aiGLM 5V Turbo1400
  47. #47AlibabaQwen3.5 397B A17B1399
  48. #48OpenAIGPT-5.4 Mini1397
  49. #48MiniMaxMiniMax M2.71397
  50. #50MiniMaxMiniMax M2.51386
  51. #51AnthropicClaude Sonnet 4.51384
  52. #52xAIGrok 4.20 (Reasoning)1374
  53. #53UpstageSolar Pro 41370
  54. #54GoogleGemma 4 31B1365
  55. #55xAIGrok 4.31357
  56. #56TencentHy31356
  57. #57AnthropicClaude Haiku 4.51329
  58. #58DeepSeekDeepSeek V3.21325
  59. #59Mistral AIMistral Medium 3.51263
  60. #60GoogleGemini 3.1 Flash Lite1256
  61. #61xAIGrok 4.1 Fast (Reasoning)1241
  62. #62Arcee AITrinity Large Thinking1238
  63. #63GoogleGemini 2.5 Pro1227
  64. #64IBMGranite 4.1 8B1190
Half of models ≤ 1486Measured by: User preference votesLast updated 2026-10-01

18 of the 34 AIs released in the last 60 days have no score here yet · Mistral Large 4, Clef, Perplexity Decider v1 27B and 15 more

Source: Arena Intelligence Data sources & removal requests

What to keep in mind

  • It’s a preference vote, not a check that the request was followed. A result that looks or sounds nicer can win even if it strays.
  • Which requests got made depends on who voted — they may differ from your use.
  • Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.