Arena Mathematical

In plain words · Which AI people preferred on expert-level mathematics questions

What does it measure?
Which AI's answers were preferred on expert-level questions in Mathematical & Statistics. Examples: checking proofs, choosing statistical models, optimisation problems. These are questions the people who actually do that work would ask, not general ones.
Who evaluates it, and how?
On Arena Intelligence, a user enters a question, two anonymous AIs answer, and the user votes for the better one. Over 82 million votes are aggregated with a statistical model (Bradley-Terry) into an Elo-style score, and this table takes the 'style-controlled' version from the official dataset, which statistically removes the bias toward long, nicely formatted answers. An AI classifier judges each prompt's reasoning depth and expertise and tags only about 5.5% as 'expert' prompts, which are then divided into occupational categories based on the US Bureau of Labor Statistics classification (SOC), launched November 2025. This table counts only prompts classified as 'Mathematical & Statistics'.
Reading the score
This is a head-to-head record, not an exam score (accuracy %). It is computed from the probability of being preferred over other AIs, so there is no maximum, and when a new AI enters, existing scores shift slightly. Typical values are 1,000–1,500; a 100-point gap means the leader is preferred about 64% of the time, a 200-point gap about 76%. Correctness is not counted — a wrong but friendly answer can win — so if accuracy matters, also look at accuracy tests (GPQA, AIME, etc.). Gaps of a few dozen points are effectively the same tier even if the rank differs. This category has few prompts, so top ranks change often; look at score bands rather than ranks.

Last updated: 2026-10-02

Rank

Top score1542Gemini 4 Argon
Models tested94
Last updated2026.10.02
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 🥇GoogleGemini 4 Argon1542
  2. 🥈AnthropicClaude Fable 5.11523
  3. 🥉AnthropicClaude Fable 51522
  4. #4AnthropicClaude Opus 5.51521
  5. #5AnthropicClaude Opus 51518
  6. #6AnthropicClaude Opus 4.61513
  7. #7OpenAIGPT-5.6 Sol1513
  8. #8GoogleGemini 3.8 Flash1512
  9. #9Moonshot AIKimi K31509
  10. #10DeepSeekDeepSeek V4.1 Flash1503
  11. #11XiaomiMiMo-V2.6-Pro1501
  12. #12MetaMuse Spark 1.11500
  13. #13Z.aiGLM 5.31499
  14. #14OpenAIGPT-5.41499
  15. #15OpenAIGPT-5.51499
  16. #16AnthropicClaude Opus 4.71499
  17. #17OpenAIGPT-6 Astra1498
  18. #18GoogleGemini 3.7 Flash1495
  19. #19Z.aiGLM 5.3 Flash1495
  20. #20MetaMuse Spark 1.21494
  21. #21Z.aiGLM 5.21494
  22. #22xAIGrok 4.51493
  23. #23MetaMuse Spark 1.31493
  24. #24XiaomiMiMo V2.5 Pro1491
  25. #25AnthropicClaude Opus 4.81490
  26. #26OpenAIGPT-5.6 Terra1489
  27. #27GoogleGemini 3.5 Flash1488
  28. #28GoogleGemini 3.6 Flash1487
  29. #29AlibabaQwen3.6 Max1487
  30. #30AnthropicClaude Sonnet 51487
  31. #31OpenAIGPT-5.6 Luna1487
  32. #32xAIGrok 4.61485
  33. #33Moonshot AIKimi K2.61485
  34. #34AnthropicClaude Sonnet 4.61485
  35. #35GoogleGemini 3.1 Pro1484
  36. #36Moonshot AIKimi K2.51477
  37. #37AnthropicClaude Opus 4.51475
  38. #38Z.aiGLM-5.11474
  39. #39AlibabaQwen3.7 Plus1474
  40. #40GoogleGemma 4 31B1473
  41. #41XiaomiMiMo V2 Pro1471
  42. #41MetaMuse Spark1471
  43. #43GoogleGemini 3 Flash1468
  44. #44OpenAIGPT-6 Luna1468
  45. #45AlibabaQwen3.7 Max1466
  46. #46xAIGrok 4.20 (Reasoning)1460
  47. #47DeepSeekDeepSeek V4 Pro1460
  48. #48Z.aiGLM-51458
  49. #49XiaomiMiMo-V2.6-Flash1457
  50. #50AlibabaQwen3.6 Plus1455
  51. #51xAIGrok 4.201455
  52. #52AlibabaQwen3.5 397B A17B1454
  53. #53XiaomiMiMo V2.51454
  54. #54OpenAIGPT-6 Sol1454
  55. #55AnthropicClaude Sonnet 4.51452
  56. #55OpenAIGPT-5.4 Mini1452
  57. #55TencentHy31452
  58. #58NVIDIANemotron 3 Ultra1451
  59. #59AnthropicClaude Opus 4.11451
  60. #60xAIGrok 4.71450
  61. #61ByteDanceDola Seed 2.0 Pro1448
  62. #62GoogleGemini 3.5 Flash-Lite1448
  63. #62MiniMaxMiniMax M31448
  64. #64GoogleGemini 2.5 Pro1447
  65. #65Z.aiGLM 5V Turbo1446
  66. #66OpenAIGPT-51442
  67. #67MiniMaxMiniMax M2.71441
  68. #68MeituanLongcat Flash Chat1440
  69. #69Mistral AIMistral Medium 3.51439
  70. #70BaiduERNIE 5.0 Thinking1435
  71. #71DeepSeekDeepSeek V4 Flash1434
  72. #72AnthropicClaude Haiku 4.51433
  73. #73DeepSeekDeepSeek V3.21432
  74. #74GoogleGemini 3.1 Flash Lite1432
  75. #75xAIGrok 4.31431
  76. #76OpenAIGPT-5.4 Nano1431
  77. #77AnthropicClaude Opus 41426
  78. #78xAIGrok 4.1 Fast (Reasoning)1423
  79. #79UpstageSolar Pro 41421
  80. #80GoogleGemini 2.5 Flash1416
  81. #81AnthropicClaude Sonnet 41413
  82. #82OpenAIGPT-5 Mini1407
  83. #83MiniMaxMiniMax M2.51399
  84. #83NVIDIANemotron 3 Super1399
  85. #85OpenAIGPT OSS 120B1384
  86. #86Arcee AITrinity Large Thinking1376
  87. #87GoogleGemini 2.5 Flash Lite1370
  88. #88OpenAIGPT-5 Nano1355
  89. #89NVIDIANemotron 3 Nano 30B A3B1354
  90. #90AmazonNova 2 Lite1347
  91. #91IBMGranite 4.1 8B1326
  92. #92MetaLlama 4 Maverick1319
  93. #93MetaLlama 4 Scout1314
  94. #94OpenAIGPT-4.11309
Half of models ≤ 1459Measured by: User preference votesLast updated 2026-10-02

21 of the 37 AIs released in the last 60 days have no score here yet · Claude Haiku 5.5, Decider V1.1 27B, GPT-6 Luna Decisions and 18 more

Source: Arena Intelligence Data sources & removal requests

What to keep in mind

  • It’s a preference vote, not a correctness check. A wrong but friendly answer can win.
  • Which questions got asked depends on who voted — they may differ from your use.
  • Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.