Arena Search
Often citedHigher is better
Users ask a question, two anonymous search-enabled AIs answer with citations, and the user votes for the better answer. Citation formatting is standardized so the models cannot be recognized by style. Scores are Elo ratings from head-to-head human votes. Higher is better.
Top score1257GPT-5.6 Sol
Models tested7
Last updated2026.08.24
Model release date (newest on the right)
Best score so farHigher on the chart is better
Which AI gives answers people prefer most, without false claims?Arena Text Factuality
Gemini 4 Argon🥇
Gemini 4 Argon🥇Does it know facts accurately without searching?SimpleQA Verified
GPT-6 Astra🥇
GPT-6 Astra🥇Can it carry out multi-step web research?Deep Research Bench
Claude Opus 4.6🥇
Claude Opus 4.6🥇
- 🥇OpenAIGPT-5.6 Sol1257
- 🥈AnthropicClaude Opus 4.71233
- 🥉AnthropicClaude Fable 51230
- #4xAIGrok 4.51213
- #5AnthropicClaude Opus 4.81204
- #6xAIGrok 4.201189
- #7xAIGrok 4.31165
Half of models ≤ 1213Last updated 2026-10-07