Arena Search

Often citedHigher is better

Users ask a question, two anonymous search-enabled AIs answer with citations, and the user votes for the better answer. Citation formatting is standardized so the models cannot be recognized by style. Scores are Elo ratings from head-to-head human votes. Higher is better.

Top score1257GPT-5.6 Sol
Models tested7
Last updated2026.08.24
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 🥇OpenAIGPT-5.6 Sol1257
  2. 🥈AnthropicClaude Opus 4.71233
  3. 🥉AnthropicClaude Fable 51230
  4. #4xAIGrok 4.51213
  5. #5AnthropicClaude Opus 4.81204
  6. #6xAIGrok 4.201189
  7. #7xAIGrok 4.31165
Half of models ≤ 1213Last updated 2026-10-07