Arena Entertainment, Sports & Media
In plain words · Which AI people preferred on expert-level entertainment, sports and media questions
- What does it measure?
- Which AI's answers were preferred on expert-level questions in Entertainment, Sports & Media. Examples: script structure, content planning, sports analysis. These are questions the people who actually do that work would ask, not general ones.
- Who evaluates it, and how?
- On Arena Intelligence, a user enters a question, two anonymous AIs answer, and the user votes for the better one. Over 82 million votes are aggregated with a statistical model (Bradley-Terry) into an Elo-style score, and this table takes the 'style-controlled' version from the official dataset, which statistically removes the bias toward long, nicely formatted answers. An AI classifier judges each prompt's reasoning depth and expertise and tags only about 5.5% as 'expert' prompts, which are then divided into occupational categories based on the US Bureau of Labor Statistics classification (SOC), launched November 2025. This table counts only prompts classified as 'Entertainment, Sports & Media'.
- Reading the score
- This is a head-to-head record, not an exam score (accuracy %). It is computed from the probability of being preferred over other AIs, so there is no maximum, and when a new AI enters, existing scores shift slightly. Typical values are 1,000–1,500; a 100-point gap means the leader is preferred about 64% of the time, a 200-point gap about 76%. Correctness is not counted — a wrong but friendly answer can win — so if accuracy matters, also look at accuracy tests (GPQA, AIME, etc.). Gaps of a few dozen points are effectively the same tier even if the rank differs. This category has few prompts, so top ranks change often; look at score bands rather than ranks.
Last updated: 2026-10-02
Rank
Top score1522Gemini 4 Argon
Models tested97
Last updated2026.10.02
Model release date (newest on the right)
Best score so farHigher on the chart is better
Which AI gives the writing and literature answers people prefer most?Arena Writing, Literature & Language
Gemini 4 Argon🥇
Gemini 4 Argon🥇Which AI do people prefer most in long conversations?Arena Multi-turn
Gemini 4 Argon🥇
Gemini 4 Argon🥇Does it follow tricky rules like format and length without missing any?IFBench
Grok 4.3🥇
Grok 4.3🥇
- 🥇GoogleGemini 4 Argon1522
- 🥈AnthropicClaude Opus 5.51501
- 🥉AnthropicClaude Fable 51495
- #4GoogleGemini 3.7 Flash1480
- #5AnthropicClaude Opus 4.61478
- #6AnthropicClaude Fable 5.11475
- #7AnthropicClaude Opus 4.71473
- #8GoogleGemini 3.8 Flash1471
- #9OpenAIGPT-5.6 Sol1471
- #10GoogleGemini 3.1 Pro1469
- #11MetaMuse Spark1462
- #12OpenAIGPT-6.1 Sol1462
- #13GoogleGemini 3.6 Flash1461
- #14AnthropicClaude Opus 51461
- #15MetaMuse Spark 1.21460
- #16Moonshot AIKimi K31460
- #17OpenAIGPT-6 Astra1458
- #18GoogleGemini 3.5 Flash1457
- #19MetaMuse Spark 1.31456
- #20AnthropicClaude Opus 4.51455
- #21AnthropicClaude Opus 4.81454
- #22GoogleGemini 3 Flash1453
- #23DeepSeekDeepSeek V4.1 Flash1453
- #24MetaMuse Spark 1.11450
- #25xAIGrok 4.201449
- #26OpenAIGPT-5.51447
- #26XiaomiMiMo-V2.6-Pro1447
- #28Z.aiGLM 5.31447
- #29Z.aiGLM 5.21446
- #30AnthropicClaude Sonnet 4.61446
- #31xAIGrok 4.51445
- #32xAIGrok 4.20 (Reasoning)1445
- #33AnthropicClaude Sonnet 4.51444
- #34AlibabaQwen3.7 Max1443
- #35xAIGrok 4.61443
- #36Z.aiGLM-5.11442
- #37OpenAIGPT-6 Sol1442
- #38OpenAIGPT-5.41441
- #39AnthropicClaude Sonnet 5.51441
- #40Z.aiGLM-51439
- #41Z.aiGLM 5.3 Flash1439
- #42XiaomiMiMo V2.5 Pro1438
- #43AnthropicClaude Opus 4.11435
- #44DeepSeekDeepSeek V4 Pro1435
- #45GoogleGemini 3.5 Flash-Lite1433
- #46AnthropicClaude Sonnet 51433
- #47OpenAIGPT-5.6 Terra1432
- #48GoogleGemini 2.5 Pro1431
- #49Moonshot AIKimi K2.61429
- #50AlibabaQwen3.6 Max1427
- #51XiaomiMiMo V2 Pro1426
- #52OpenAIGPT-6 Luna1422
- #53AnthropicClaude Opus 41422
- #54Moonshot AIKimi K2.51421
- #55StepfunStep-51420
- #56AlibabaQwen3.7 Plus1419
- #57OpenAIGPT-5.6 Luna1418
- #58GoogleGemma 4 31B1417
- #59xAIGrok 4.71416
- #60xAIGrok 4.31416
- #61ByteDanceDola Seed 2.0 Pro1414
- #62BaiduERNIE 5.0 Thinking1412
- #63OpenAIGPT-5.4 Mini1409
- #64AlibabaQwen3.6 Plus1407
- #65XiaomiMiMo-V2.6-Flash1407
- #66MiniMaxMiniMax M31407
- #67DeepSeekDeepSeek V4 Flash1405
- #68GoogleGemini 3.1 Flash Lite1404
- #69XiaomiMiMo V2.51403
- #70xAIGrok 4.1 Fast (Reasoning)1402
- #71AlibabaQwen3.5 397B A17B1402
- #72OpenAIGPT-51398
- #73Z.aiGLM 5V Turbo1398
- #74MeituanLongcat Flash Chat1397
- #75Mistral AIMistral Medium 3.51396
- #76DeepSeekDeepSeek V3.21396
- #77AnthropicClaude Sonnet 41389
- #78AnthropicClaude Haiku 4.51389
- #79GoogleGemini 2.5 Flash1386
- #80NVIDIANemotron 3 Ultra1378
- #81MiniMaxMiniMax M2.71372
- #82TencentHy31364
- #83MiniMaxMiniMax M2.51362
- #84OpenAIGPT-5.4 Nano1352
- #85OpenAIGPT-5 Mini1348
- #86GoogleGemini 2.5 Flash Lite1346
- #87UpstageSolar Pro 41326
- #88Arcee AITrinity Large Thinking1324
- #89NVIDIANemotron 3 Super1316
- #90MetaLlama 4 Maverick1296
- #91OpenAIGPT-4.11291
- #92MetaLlama 4 Scout1289
- #93OpenAIGPT OSS 120B1286
- #94AmazonNova 2 Lite1283
- #95OpenAIGPT-5 Nano1277
- #96NVIDIANemotron 3 Nano 30B A3B1257
- #97IBMGranite 4.1 8B1256
Half of models ≤ 1429Measured by: User preference votesLast updated 2026-10-02
18 of the 37 AIs released in the last 60 days have no score here yet · Claude Haiku 5.5, Decider V1.1 27B, GPT-6 Luna Decisions and 15 more
Source: Arena Intelligence Data sources & removal requestsWhat to keep in mind
- It’s a preference vote, not a correctness check. A wrong but friendly answer can win.
- Which questions got asked depends on who voted — they may differ from your use.
- Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.