Arena Document

Higher is better

Users upload a PDF and ask for answers, summaries, or extracted information; two anonymous models respond and the user votes for the better one. It reflects how well models read and analyze long, real user documents. Scores are Elo ratings from head-to-head human votes. Higher is better.

Top score1516Claude Opus 5
Models tested34
Last updated2026.09.13
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 🥇AnthropicClaude Opus 5Reasoning effort: High1516
  2. 🥈AnthropicClaude Fable 5.11513
  3. 🥉AnthropicClaude Opus 4.6Reasoning effort: High1507
  4. #4AnthropicClaude Fable 51496
  5. #5AnthropicClaude Opus 4.71495
  6. #6OpenAIGPT-5.51486
  7. #7OpenAIGPT-5.6 Sol1483
  8. #8AnthropicClaude Sonnet 4.61482
  9. #9AnthropicClaude Opus 4.8Reasoning effort: High1475
  10. #10OpenAIGPT-5.6 Terra1472
  11. #11OpenAIGPT-5.41471
  12. #12MetaMuse Spark 1.31471
  13. #13OpenAIGPT-6 Astra1468
  14. #14AnthropicClaude Sonnet 51466
  15. #15MetaMuse Spark 1.11465
  16. #16GoogleGemini 3.5 Flash1463
  17. #17AnthropicClaude Opus 4.51462
  18. #18OpenAIGPT-5.6 Luna1457
  19. #19GoogleGemini 3.6 Flash1456
  20. #20xAIGrok 4.61452
  21. #21xAIGrok 4.51452
  22. #22Moonshot AIKimi K2.61451
  23. #23AnthropicClaude Sonnet 4.51450
  24. #24MetaMuse Spark1444
  25. #25AlibabaQwen3.7 Plus1444
  26. #26GoogleGemini 3.1 Pro1444
  27. #27MiniMaxMiniMax M31435
  28. #28Moonshot AIKimi K2.51430
  29. #29GoogleGemma 4 31B1425
  30. #30GoogleGemini 2.5 Pro1421
  31. #31AnthropicClaude Haiku 4.51420
  32. #32xAIGrok 4.20 (Reasoning)1416
  33. #33Z.aiGLM 5V Turbo1416
  34. #34GoogleGemini 3 Flash1413
Top 3 within the margin of error · effectively tied · Half of models ≤ 1460Last updated 2026-10-07

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.