Arena Document
Higher is better
Users upload a PDF and ask for answers, summaries, or extracted information; two anonymous models respond and the user votes for the better one. It reflects how well models read and analyze long, real user documents. Scores are Elo ratings from head-to-head human votes. Higher is better.
Top score1516Claude Opus 5
Models tested34
Last updated2026.09.13
Model release date (newest on the right)
Best score so farHigher on the chart is better
Can it pull scattered facts from several long documents into one answer?Long Context Reasoning
Kimi K3🥇
Kimi K3🥇Can it read tables and reshape them as needed?LiveBench Data Analysis
GPT-6 Astra🥇
GPT-6 Astra🥇Which AI gives the business answers people prefer most?Arena Business, Management & Finance
Gemini 4 Argon🥇
Gemini 4 Argon🥇
- 🥇AnthropicClaude Opus 5Reasoning effort: High1516
- 🥈AnthropicClaude Fable 5.11513
- 🥉AnthropicClaude Opus 4.6Reasoning effort: High1507
- #4AnthropicClaude Fable 51496
- #5AnthropicClaude Opus 4.71495
- #6OpenAIGPT-5.51486
- #7OpenAIGPT-5.6 Sol1483
- #8AnthropicClaude Sonnet 4.61482
- #9AnthropicClaude Opus 4.8Reasoning effort: High1475
- #10OpenAIGPT-5.6 Terra1472
- #11OpenAIGPT-5.41471
- #12MetaMuse Spark 1.31471
- #13OpenAIGPT-6 Astra1468
- #14AnthropicClaude Sonnet 51466
- #15MetaMuse Spark 1.11465
- #16GoogleGemini 3.5 Flash1463
- #17AnthropicClaude Opus 4.51462
- #18OpenAIGPT-5.6 Luna1457
- #19GoogleGemini 3.6 Flash1456
- #20xAIGrok 4.61452
- #21xAIGrok 4.51452
- #22Moonshot AIKimi K2.61451
- #23AnthropicClaude Sonnet 4.51450
- #24MetaMuse Spark1444
- #25AlibabaQwen3.7 Plus1444
- #26GoogleGemini 3.1 Pro1444
- #27MiniMaxMiniMax M31435
- #28Moonshot AIKimi K2.51430
- #29GoogleGemma 4 31B1425
- #30GoogleGemini 2.5 Pro1421
- #31AnthropicClaude Haiku 4.51420
- #32xAIGrok 4.20 (Reasoning)1416
- #33Z.aiGLM 5V Turbo1416
- #34GoogleGemini 3 Flash1413
Top 3 within the margin of error · effectively tied · Half of models ≤ 1460Last updated 2026-10-07
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.