Arena Business, Management & Finance

In plain words · Which AI people preferred on expert-level business, management and finance questions

What does it measure?
Which AI's answers were preferred on expert-level questions in Business, Management & Finance. Examples: reviewing financial models, business plans, accounting treatment. These are questions the people who actually do that work would ask, not general ones.
Who evaluates it, and how?
On Arena Intelligence, a user enters a question, two anonymous AIs answer, and the user votes for the better one. Over 82 million votes are aggregated with a statistical model (Bradley-Terry) into an Elo-style score, and this table takes the 'style-controlled' version from the official dataset, which statistically removes the bias toward long, nicely formatted answers. An AI classifier judges each prompt's reasoning depth and expertise and tags only about 5.5% as 'expert' prompts, which are then divided into occupational categories based on the US Bureau of Labor Statistics classification (SOC), launched November 2025. This table counts only prompts classified as 'Business, Management & Finance'.
Reading the score
This is a head-to-head record, not an exam score (accuracy %). It is computed from the probability of being preferred over other AIs, so there is no maximum, and when a new AI enters, existing scores shift slightly. Typical values are 1,000–1,500; a 100-point gap means the leader is preferred about 64% of the time, a 200-point gap about 76%. Correctness is not counted — a wrong but friendly answer can win — so if accuracy matters, also look at accuracy tests (GPQA, AIME, etc.). Gaps of a few dozen points are effectively the same tier even if the rank differs. This category has few prompts, so top ranks change often; look at score bands rather than ranks.

Last updated: 2026-10-02

Rank

Top score1523Gemini 4 Argon
Models tested97
Last updated2026.10.02
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 🥇GoogleGemini 4 Argon79.7%—1523
  2. 🥈MetaMuse Spark 1.279.0%76.51514
  3. 🥉MetaMuse Spark 1.383.0%79.61505
  4. #4AnthropicClaude Fable 582.3%80.51502
  5. #5AnthropicClaude Opus 4.678.0%69.91500
  6. #6AnthropicClaude Opus 4.778.7%78.31496
  7. #7MetaMuse Spark 1.177.7%72.51493
  8. #8MetaMuse Spark78.0%—1489
  9. #9AnthropicClaude Opus 5.584.7%79.81489
  10. #10Moonshot AIKimi K388.7%78.71487
  11. #11AnthropicClaude Opus 4.877.7%66.01487
  12. #12OpenAIGPT-5.584.3%81.61485
  13. #13OpenAIGPT-5.6 Sol84.0%79.81485
  14. #14AnthropicClaude Opus 582.0%74.51484
  15. #15AnthropicClaude Fable 5.185.3%80.31484
  16. #16GoogleGemini 3.7 Flash83.0%68.01483
  17. #17OpenAIGPT-5.482.0%79.31482
  18. #17AnthropicClaude Sonnet 4.680.0%78.01482
  19. #19GoogleGemini 3.8 Flash84.0%54.01481
  20. #20DeepSeekDeepSeek V4.1 Flash84.0%79.31479
  21. #21GoogleGemini 3.1 Pro82.0%78.51477
  22. #22Z.aiGLM 5.379.7%70.21475
  23. #23OpenAIGPT-6.1 Sol84.0%82.21475
  24. #24OpenAIGPT-6 Astra80.7%83.01473
  25. #25GoogleGemini 3.6 Flash80.0%63.01472
  26. #26AnthropicClaude Opus 4.577.3%74.41472
  27. #27XiaomiMiMo V2.5 Pro79.7%—1471
  28. #28xAIGrok 4.2023.3%—1470
  29. #29GoogleGemini 3 Flash78.0%—1468
  30. #30StepfunStep-588.3%—1468
  31. #31OpenAIGPT-5.6 Terra83.0%79.31467
  32. #31Z.aiGLM 5.3 Flash80.0%76.41467
  33. #33xAIGrok 4.20 (Reasoning)69.0%—1466
  34. #34Z.aiGLM 5.278.3%73.71466
  35. #35xAIGrok 4.579.3%73.01465
  36. #36Z.aiGLM-5.173.7%—1465
  37. #37AnthropicClaude Sonnet 582.0%71.71465
  38. #38ByteDanceDola Seed 2.0 Pro——1464
  39. #39AnthropicClaude Sonnet 5.582.7%78.61464
  40. #40OpenAIGPT-5.6 Luna83.7%78.01463
  41. #41GoogleGemini 3.5 Flash74.3%64.91462
  42. #42XiaomiMiMo-V2.6-Pro86.3%—1461
  43. #43AnthropicClaude Sonnet 4.572.3%—1461
  44. #44DeepSeekDeepSeek V4 Pro80.3%74.51460
  45. #45OpenAIGPT-5.4 Mini77.0%70.81460
  46. #46XiaomiMiMo-V2.6-Flash74.3%—1459
  47. #47AlibabaQwen3.7 Max79.0%71.81458
  48. #48Moonshot AIKimi K2.681.0%65.11455
  49. #49XiaomiMiMo V2 Pro68.3%—1454
  50. #50GoogleGemini 3.5 Flash-Lite76.0%53.31452
  51. #51AlibabaQwen3.6 Max80.7%—1451
  52. #52Z.aiGLM-575.7%—1451
  53. #53GoogleGemma 4 31B69.7%—1449
  54. #54MiniMaxMiniMax M383.0%76.21448
  55. #55AnthropicClaude Opus 4.176.0%—1447
  56. #56OpenAIGPT-6 Luna83.3%73.41447
  57. #56xAIGrok 4.681.0%73.91447
  58. #58AlibabaQwen3.5 397B A17B77.3%—1446
  59. #59OpenAIGPT-6 Sol83.7%81.21446
  60. #60AlibabaQwen3.6 Plus78.3%69.91446
  61. #61AlibabaQwen3.7 Plus73.0%—1445
  62. #62XiaomiMiMo V2.573.0%—1441
  63. #63BaiduERNIE 5.0 Thinking7.0%—1441
  64. #64Moonshot AIKimi K2.578.0%—1440
  65. #65Z.aiGLM 5V Turbo70.3%—1440
  66. #66xAIGrok 4.375.0%55.81439
  67. #67DeepSeekDeepSeek V4 Flash79.7%68.01439
  68. #68GoogleGemini 2.5 Pro69.0%—1437
  69. #69xAIGrok 4.780.0%76.91433
  70. #70MeituanLongcat Flash Chat28.3%—1433
  71. #71Mistral AIMistral Medium 3.569.3%—1430
  72. #72MiniMaxMiniMax M2.778.3%—1424
  73. #73DeepSeekDeepSeek V3.273.3%—1422
  74. #74GoogleGemini 3.1 Flash Lite74.3%—1420
  75. #75AnthropicClaude Haiku 4.574.3%—1420
  76. #76NVIDIANemotron 3 Ultra79.3%54.51418
  77. #77xAIGrok 4.1 Fast (Reasoning)74.0%—1417
  78. #78OpenAIGPT-578.2%—1413
  79. #79TencentHy379.0%—1412
  80. #80AnthropicClaude Opus 469.3%—1409
  81. #81OpenAIGPT-5.4 Nano76.7%67.61405
  82. #82MiniMaxMiniMax M2.573.3%—1400
  83. #83GoogleGemini 2.5 Flash65.3%—1399
  84. #84UpstageSolar Pro 474.0%—1388
  85. #85OpenAIGPT-5 Mini72.3%—1381
  86. #86AnthropicClaude Sonnet 470.3%—1381
  87. #87GoogleGemini 2.5 Flash Lite64.7%—1378
  88. #88Arcee AITrinity Large Thinking38.0%—1367
  89. #89OpenAIGPT OSS 120B52.0%—1352
  90. #90AmazonNova 2 Lite60.3%—1351
  91. #91NVIDIANemotron 3 Super65.7%—1350
  92. #92OpenAIGPT-5 Nano45.0%—1339
  93. #93MetaLlama 4 Maverick50.0%—1317
  94. #93NVIDIANemotron 3 Nano 30B A3B38.0%—1317
  95. #95MetaLlama 4 Scout27.7%—1316
  96. #96IBMGranite 4.1 8B13.3%—1293
  97. #97OpenAIGPT-4.168.3%—1279
Half of models ≤ 1454Measured by: User preference votesLast updated 2026-10-02

18 of the 37 AIs released in the last 60 days have no score here yet · Claude Haiku 5.5, Decider V1.1 27B, GPT-6 Luna Decisions and 15 more

Source: Arena Intelligence Data sources & removal requests

What to keep in mind

  • It’s a preference vote, not a correctness check. A wrong but friendly answer can win.
  • Which questions got asked depends on who voted — they may differ from your use.
  • Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.