Arena Entertainment, Sports & Media

In plain words · Which AI people preferred on expert-level entertainment, sports and media questions

What does it measure?
Which AI's answers were preferred on expert-level questions in Entertainment, Sports & Media. Examples: script structure, content planning, sports analysis. These are questions the people who actually do that work would ask, not general ones.
Who evaluates it, and how?
On Arena Intelligence, a user enters a question, two anonymous AIs answer, and the user votes for the better one. Over 82 million votes are aggregated with a statistical model (Bradley-Terry) into an Elo-style score, and this table takes the 'style-controlled' version from the official dataset, which statistically removes the bias toward long, nicely formatted answers. An AI classifier judges each prompt's reasoning depth and expertise and tags only about 5.5% as 'expert' prompts, which are then divided into occupational categories based on the US Bureau of Labor Statistics classification (SOC), launched November 2025. This table counts only prompts classified as 'Entertainment, Sports & Media'.
Reading the score
This is a head-to-head record, not an exam score (accuracy %). It is computed from the probability of being preferred over other AIs, so there is no maximum, and when a new AI enters, existing scores shift slightly. Typical values are 1,000–1,500; a 100-point gap means the leader is preferred about 64% of the time, a 200-point gap about 76%. Correctness is not counted — a wrong but friendly answer can win — so if accuracy matters, also look at accuracy tests (GPQA, AIME, etc.). Gaps of a few dozen points are effectively the same tier even if the rank differs. This category has few prompts, so top ranks change often; look at score bands rather than ranks.

Last updated: 2026-10-02

Rank

Top score1522Gemini 4 Argon
Models tested97
Last updated2026.10.02
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 🥇GoogleGemini 4 Argon1522
  2. 🥈AnthropicClaude Opus 5.51501
  3. 🥉AnthropicClaude Fable 51495
  4. #4GoogleGemini 3.7 Flash1480
  5. #5AnthropicClaude Opus 4.61478
  6. #6AnthropicClaude Fable 5.11475
  7. #7AnthropicClaude Opus 4.71473
  8. #8GoogleGemini 3.8 Flash1471
  9. #9OpenAIGPT-5.6 Sol1471
  10. #10GoogleGemini 3.1 Pro1469
  11. #11MetaMuse Spark1462
  12. #12OpenAIGPT-6.1 Sol1462
  13. #13GoogleGemini 3.6 Flash1461
  14. #14AnthropicClaude Opus 51461
  15. #15MetaMuse Spark 1.21460
  16. #16Moonshot AIKimi K31460
  17. #17OpenAIGPT-6 Astra1458
  18. #18GoogleGemini 3.5 Flash1457
  19. #19MetaMuse Spark 1.31456
  20. #20AnthropicClaude Opus 4.51455
  21. #21AnthropicClaude Opus 4.81454
  22. #22GoogleGemini 3 Flash1453
  23. #23DeepSeekDeepSeek V4.1 Flash1453
  24. #24MetaMuse Spark 1.11450
  25. #25xAIGrok 4.201449
  26. #26OpenAIGPT-5.51447
  27. #26XiaomiMiMo-V2.6-Pro1447
  28. #28Z.aiGLM 5.31447
  29. #29Z.aiGLM 5.21446
  30. #30AnthropicClaude Sonnet 4.61446
  31. #31xAIGrok 4.51445
  32. #32xAIGrok 4.20 (Reasoning)1445
  33. #33AnthropicClaude Sonnet 4.51444
  34. #34AlibabaQwen3.7 Max1443
  35. #35xAIGrok 4.61443
  36. #36Z.aiGLM-5.11442
  37. #37OpenAIGPT-6 Sol1442
  38. #38OpenAIGPT-5.41441
  39. #39AnthropicClaude Sonnet 5.51441
  40. #40Z.aiGLM-51439
  41. #41Z.aiGLM 5.3 Flash1439
  42. #42XiaomiMiMo V2.5 Pro1438
  43. #43AnthropicClaude Opus 4.11435
  44. #44DeepSeekDeepSeek V4 Pro1435
  45. #45GoogleGemini 3.5 Flash-Lite1433
  46. #46AnthropicClaude Sonnet 51433
  47. #47OpenAIGPT-5.6 Terra1432
  48. #48GoogleGemini 2.5 Pro1431
  49. #49Moonshot AIKimi K2.61429
  50. #50AlibabaQwen3.6 Max1427
  51. #51XiaomiMiMo V2 Pro1426
  52. #52OpenAIGPT-6 Luna1422
  53. #53AnthropicClaude Opus 41422
  54. #54Moonshot AIKimi K2.51421
  55. #55StepfunStep-51420
  56. #56AlibabaQwen3.7 Plus1419
  57. #57OpenAIGPT-5.6 Luna1418
  58. #58GoogleGemma 4 31B1417
  59. #59xAIGrok 4.71416
  60. #60xAIGrok 4.31416
  61. #61ByteDanceDola Seed 2.0 Pro1414
  62. #62BaiduERNIE 5.0 Thinking1412
  63. #63OpenAIGPT-5.4 Mini1409
  64. #64AlibabaQwen3.6 Plus1407
  65. #65XiaomiMiMo-V2.6-Flash1407
  66. #66MiniMaxMiniMax M31407
  67. #67DeepSeekDeepSeek V4 Flash1405
  68. #68GoogleGemini 3.1 Flash Lite1404
  69. #69XiaomiMiMo V2.51403
  70. #70xAIGrok 4.1 Fast (Reasoning)1402
  71. #71AlibabaQwen3.5 397B A17B1402
  72. #72OpenAIGPT-51398
  73. #73Z.aiGLM 5V Turbo1398
  74. #74MeituanLongcat Flash Chat1397
  75. #75Mistral AIMistral Medium 3.51396
  76. #76DeepSeekDeepSeek V3.21396
  77. #77AnthropicClaude Sonnet 41389
  78. #78AnthropicClaude Haiku 4.51389
  79. #79GoogleGemini 2.5 Flash1386
  80. #80NVIDIANemotron 3 Ultra1378
  81. #81MiniMaxMiniMax M2.71372
  82. #82TencentHy31364
  83. #83MiniMaxMiniMax M2.51362
  84. #84OpenAIGPT-5.4 Nano1352
  85. #85OpenAIGPT-5 Mini1348
  86. #86GoogleGemini 2.5 Flash Lite1346
  87. #87UpstageSolar Pro 41326
  88. #88Arcee AITrinity Large Thinking1324
  89. #89NVIDIANemotron 3 Super1316
  90. #90MetaLlama 4 Maverick1296
  91. #91OpenAIGPT-4.11291
  92. #92MetaLlama 4 Scout1289
  93. #93OpenAIGPT OSS 120B1286
  94. #94AmazonNova 2 Lite1283
  95. #95OpenAIGPT-5 Nano1277
  96. #96NVIDIANemotron 3 Nano 30B A3B1257
  97. #97IBMGranite 4.1 8B1256
Half of models ≤ 1429Measured by: User preference votesLast updated 2026-10-02

18 of the 37 AIs released in the last 60 days have no score here yet · Claude Haiku 5.5, Decider V1.1 27B, GPT-6 Luna Decisions and 15 more

Source: Arena Intelligence Data sources & removal requests

What to keep in mind

  • It’s a preference vote, not a correctness check. A wrong but friendly answer can win.
  • Which questions got asked depends on who voted — they may differ from your use.
  • Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.