Arena Business, Management & Finance
In plain words · Which AI people preferred on expert-level business, management and finance questions
- What does it measure?
- Which AI's answers were preferred on expert-level questions in Business, Management & Finance. Examples: reviewing financial models, business plans, accounting treatment. These are questions the people who actually do that work would ask, not general ones.
- Who evaluates it, and how?
- On Arena Intelligence, a user enters a question, two anonymous AIs answer, and the user votes for the better one. Over 82 million votes are aggregated with a statistical model (Bradley-Terry) into an Elo-style score, and this table takes the 'style-controlled' version from the official dataset, which statistically removes the bias toward long, nicely formatted answers. An AI classifier judges each prompt's reasoning depth and expertise and tags only about 5.5% as 'expert' prompts, which are then divided into occupational categories based on the US Bureau of Labor Statistics classification (SOC), launched November 2025. This table counts only prompts classified as 'Business, Management & Finance'.
- Reading the score
- This is a head-to-head record, not an exam score (accuracy %). It is computed from the probability of being preferred over other AIs, so there is no maximum, and when a new AI enters, existing scores shift slightly. Typical values are 1,000–1,500; a 100-point gap means the leader is preferred about 64% of the time, a 200-point gap about 76%. Correctness is not counted — a wrong but friendly answer can win — so if accuracy matters, also look at accuracy tests (GPQA, AIME, etc.). Gaps of a few dozen points are effectively the same tier even if the rank differs. This category has few prompts, so top ranks change often; look at score bands rather than ranks.
Last updated: 2026-10-02
Rank
Top score1523Gemini 4 Argon
Models tested97
Last updated2026.10.02
Model release date (newest on the right)
Best score so farHigher on the chart is better
Can it pull scattered facts from several long documents into one answer?Long Context Reasoning
Kimi K3🥇
Kimi K3🥇Can it read tables and reshape them as needed?LiveBench Data Analysis
GPT-6 Astra🥇
GPT-6 Astra🥇
- 🥇GoogleGemini 4 Argon79.7%—1523
- 🥈MetaMuse Spark 1.279.0%76.51514
- 🥉MetaMuse Spark 1.383.0%79.61505
- #4AnthropicClaude Fable 582.3%80.51502
- #5AnthropicClaude Opus 4.678.0%69.91500
- #6AnthropicClaude Opus 4.778.7%78.31496
- #7MetaMuse Spark 1.177.7%72.51493
- #8MetaMuse Spark78.0%—1489
- #9AnthropicClaude Opus 5.584.7%79.81489
- #10Moonshot AIKimi K388.7%78.71487
- #11AnthropicClaude Opus 4.877.7%66.01487
- #12OpenAIGPT-5.584.3%81.61485
- #13OpenAIGPT-5.6 Sol84.0%79.81485
- #14AnthropicClaude Opus 582.0%74.51484
- #15AnthropicClaude Fable 5.185.3%80.31484
- #16GoogleGemini 3.7 Flash83.0%68.01483
- #17OpenAIGPT-5.482.0%79.31482
- #17AnthropicClaude Sonnet 4.680.0%78.01482
- #19GoogleGemini 3.8 Flash84.0%54.01481
- #20DeepSeekDeepSeek V4.1 Flash84.0%79.31479
- #21GoogleGemini 3.1 Pro82.0%78.51477
- #22Z.aiGLM 5.379.7%70.21475
- #23OpenAIGPT-6.1 Sol84.0%82.21475
- #24OpenAIGPT-6 Astra80.7%83.01473
- #25GoogleGemini 3.6 Flash80.0%63.01472
- #26AnthropicClaude Opus 4.577.3%74.41472
- #27XiaomiMiMo V2.5 Pro79.7%—1471
- #28xAIGrok 4.2023.3%—1470
- #29GoogleGemini 3 Flash78.0%—1468
- #30StepfunStep-588.3%—1468
- #31OpenAIGPT-5.6 Terra83.0%79.31467
- #31Z.aiGLM 5.3 Flash80.0%76.41467
- #33xAIGrok 4.20 (Reasoning)69.0%—1466
- #34Z.aiGLM 5.278.3%73.71466
- #35xAIGrok 4.579.3%73.01465
- #36Z.aiGLM-5.173.7%—1465
- #37AnthropicClaude Sonnet 582.0%71.71465
- #38ByteDanceDola Seed 2.0 Pro——1464
- #39AnthropicClaude Sonnet 5.582.7%78.61464
- #40OpenAIGPT-5.6 Luna83.7%78.01463
- #41GoogleGemini 3.5 Flash74.3%64.91462
- #42XiaomiMiMo-V2.6-Pro86.3%—1461
- #43AnthropicClaude Sonnet 4.572.3%—1461
- #44DeepSeekDeepSeek V4 Pro80.3%74.51460
- #45OpenAIGPT-5.4 Mini77.0%70.81460
- #46XiaomiMiMo-V2.6-Flash74.3%—1459
- #47AlibabaQwen3.7 Max79.0%71.81458
- #48Moonshot AIKimi K2.681.0%65.11455
- #49XiaomiMiMo V2 Pro68.3%—1454
- #50GoogleGemini 3.5 Flash-Lite76.0%53.31452
- #51AlibabaQwen3.6 Max80.7%—1451
- #52Z.aiGLM-575.7%—1451
- #53GoogleGemma 4 31B69.7%—1449
- #54MiniMaxMiniMax M383.0%76.21448
- #55AnthropicClaude Opus 4.176.0%—1447
- #56OpenAIGPT-6 Luna83.3%73.41447
- #56xAIGrok 4.681.0%73.91447
- #58AlibabaQwen3.5 397B A17B77.3%—1446
- #59OpenAIGPT-6 Sol83.7%81.21446
- #60AlibabaQwen3.6 Plus78.3%69.91446
- #61AlibabaQwen3.7 Plus73.0%—1445
- #62XiaomiMiMo V2.573.0%—1441
- #63BaiduERNIE 5.0 Thinking7.0%—1441
- #64Moonshot AIKimi K2.578.0%—1440
- #65Z.aiGLM 5V Turbo70.3%—1440
- #66xAIGrok 4.375.0%55.81439
- #67DeepSeekDeepSeek V4 Flash79.7%68.01439
- #68GoogleGemini 2.5 Pro69.0%—1437
- #69xAIGrok 4.780.0%76.91433
- #70MeituanLongcat Flash Chat28.3%—1433
- #71Mistral AIMistral Medium 3.569.3%—1430
- #72MiniMaxMiniMax M2.778.3%—1424
- #73DeepSeekDeepSeek V3.273.3%—1422
- #74GoogleGemini 3.1 Flash Lite74.3%—1420
- #75AnthropicClaude Haiku 4.574.3%—1420
- #76NVIDIANemotron 3 Ultra79.3%54.51418
- #77xAIGrok 4.1 Fast (Reasoning)74.0%—1417
- #78OpenAIGPT-578.2%—1413
- #79TencentHy379.0%—1412
- #80AnthropicClaude Opus 469.3%—1409
- #81OpenAIGPT-5.4 Nano76.7%67.61405
- #82MiniMaxMiniMax M2.573.3%—1400
- #83GoogleGemini 2.5 Flash65.3%—1399
- #84UpstageSolar Pro 474.0%—1388
- #85OpenAIGPT-5 Mini72.3%—1381
- #86AnthropicClaude Sonnet 470.3%—1381
- #87GoogleGemini 2.5 Flash Lite64.7%—1378
- #88Arcee AITrinity Large Thinking38.0%—1367
- #89OpenAIGPT OSS 120B52.0%—1352
- #90AmazonNova 2 Lite60.3%—1351
- #91NVIDIANemotron 3 Super65.7%—1350
- #92OpenAIGPT-5 Nano45.0%—1339
- #93MetaLlama 4 Maverick50.0%—1317
- #93NVIDIANemotron 3 Nano 30B A3B38.0%—1317
- #95MetaLlama 4 Scout27.7%—1316
- #96IBMGranite 4.1 8B13.3%—1293
- #97OpenAIGPT-4.168.3%—1279
Half of models ≤ 1454Measured by: User preference votesLast updated 2026-10-02
18 of the 37 AIs released in the last 60 days have no score here yet · Claude Haiku 5.5, Decider V1.1 27B, GPT-6 Luna Decisions and 15 more
Source: Arena Intelligence Data sources & removal requestsWhat to keep in mind
- It’s a preference vote, not a correctness check. A wrong but friendly answer can win.
- Which questions got asked depends on who voted — they may differ from your use.
- Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.