Arena Math

In plain words · Which AI people preferred for math questions

What does it measure?
Preference on maths questions (calculation, worked solutions, proofs). Because people choose rather than checking against an answer key, solutions that are easy to read have an advantage.
Who evaluates it, and how?
On Arena Intelligence, a user enters a question, two anonymous AIs answer, and the user votes for the better one. Over 82 million votes are aggregated with a statistical model (Bradley-Terry) into an Elo-style score, and this table takes the 'style-controlled' version from the official dataset, which statistically removes the bias toward long, nicely formatted answers. An AI classifier selects prompts that actively apply mathematical concepts — calculations and derivations.
Reading the score
This is a head-to-head record, not an exam score (accuracy %). It is computed from the probability of being preferred over other AIs, so there is no maximum, and when a new AI enters, existing scores shift slightly. Typical values are 1,000–1,500; a 100-point gap means the leader is preferred about 64% of the time, a 200-point gap about 76%. Correctness is not counted — a wrong but friendly answer can win — so if accuracy matters, also look at accuracy tests (GPQA, AIME, etc.). Gaps of a few dozen points are effectively the same tier even if the rank differs. It measures something different from accuracy tests (AIME, MATH-500): not whether the answer was right, but whether people liked it.

Last updated: 2026-10-02

Rank

Top score1530Gemini 4 Argon
Models tested94
Last updated2026.10.02
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 🥇GoogleGemini 4 Argon57.1%27.1%1530
  2. 🥈AnthropicClaude Fable 555.5%28.6%1522
  3. 🥉AnthropicClaude Fable 5.159.1%31.1%1518
  4. #4GoogleGemini 3.8 Flash47.8%18.3%1518
  5. #5AnthropicClaude Opus 554.9%29.1%1516
  6. #6AnthropicClaude Opus 5.561.4%31.7%1511
  7. #7MetaMuse Spark 1.348.7%26.0%1509
  8. #8AnthropicClaude Opus 4.639.9%—1507
  9. #9GoogleGemini 3.7 Flash47.9%14.3%1503
  10. #10Z.aiGLM 5.3 Flash39.9%15.4%1503
  11. #11DeepSeekDeepSeek V4.1 Flash39.2%14.3%1503
  12. #12GoogleGemini 3.6 Flash40.8%10.6%1501
  13. #13Moonshot AIKimi K346.9%23.4%1501
  14. #14OpenAIGPT-5.545.8%27.1%1500
  15. #15Z.aiGLM 5.342.3%19.1%1494
  16. #16OpenAIGPT-5.6 Sol49.5%32.3%1494
  17. #17MetaMuse Spark 1.146.2%15.1%1492
  18. #18OpenAIGPT-5.443.7%23.4%1492
  19. #19AnthropicClaude Opus 4.742.3%12.0%1492
  20. #20OpenAIGPT-6 Astra54.7%31.7%1490
  21. #21GoogleGemini 3.1 Pro47.0%17.7%1488
  22. #21Z.aiGLM 5.241.1%20.9%1488
  23. #23XiaomiMiMo-V2.6-Pro49.4%26.6%1487
  24. #24AlibabaQwen3.7 Max40.5%13.4%1484
  25. #25GoogleGemini 3.5 Flash42.7%13.1%1482
  26. #26XiaomiMiMo V2.5 Pro35.7%4.0%1481
  27. #27Moonshot AIKimi K2.637.5%8.0%1480
  28. #28OpenAIGPT-5.6 Terra42.9%30.0%1479
  29. #29OpenAIGPT-5.6 Luna39.5%20.6%1477
  30. #30AnthropicClaude Opus 4.848.7%20.9%1477
  31. #31AlibabaQwen3.6 Max30.8%—1476
  32. #32Z.aiGLM-5.130.1%4.6%1475
  33. #33GoogleGemini 3 Flash36.6%—1474
  34. #34AnthropicClaude Sonnet 541.3%16.9%1474
  35. #35xAIGrok 4.542.7%15.4%1473
  36. #36XiaomiMiMo-V2.6-Flash35.1%12.0%1472
  37. #37Moonshot AIKimi K2.530.7%3.1%1471
  38. #38GoogleGemma 4 31B23.6%1.4%1469
  39. #39MetaMuse Spark 1.245.5%17.7%1468
  40. #40MetaMuse Spark40.7%11.3%1466
  41. #41AlibabaQwen3.7 Plus35.6%9.1%1466
  42. #42xAIGrok 4.20 (Reasoning)34.5%—1465
  43. #43AnthropicClaude Sonnet 4.633.6%3.1%1464
  44. #44AnthropicClaude Opus 4.530.1%—1463
  45. #45OpenAIGPT-6 Luna38.5%19.4%1460
  46. #46AlibabaQwen3.6 Plus27.8%2.9%1454
  47. #47AlibabaQwen3.5 397B A17B29.0%—1452
  48. #48OpenAIGPT-6 Sol47.9%30.9%1451
  49. #49XiaomiMiMo V2 Pro30.4%—1451
  50. #50ByteDanceDola Seed 2.0 Pro——1450
  51. #51xAIGrok 4.2027.9%—1449
  52. #52xAIGrok 4.644.1%19.7%1448
  53. #53DeepSeekDeepSeek V4 Pro41.0%18.0%1445
  54. #54xAIGrok 4.743.1%18.0%1444
  55. #55Z.aiGLM-529.3%—1444
  56. #56AnthropicClaude Opus 4.112.5%—1443
  57. #57GoogleGemini 3.5 Flash-Lite18.8%0.0%1442
  58. #58NVIDIANemotron 3 Ultra28.4%3.1%1442
  59. #59GoogleGemini 2.5 Pro22.5%2.0%1439
  60. #60OpenAIGPT-5.4 Mini28.1%10.0%1439
  61. #61XiaomiMiMo V2.527.2%3.7%1438
  62. #62Z.aiGLM 5V Turbo17.1%—1437
  63. #63GoogleGemini 3.1 Flash Lite17.2%1.1%1436
  64. #64MeituanLongcat Flash Chat5.8%—1435
  65. #65BaiduERNIE 5.0 Thinking13.3%—1435
  66. #66OpenAIGPT-528.5%12.6%1434
  67. #67MiniMaxMiniMax M339.0%3.7%1432
  68. #68TencentHy333.5%—1431
  69. #69Mistral AIMistral Medium 3.513.8%0.0%1429
  70. #70AnthropicClaude Sonnet 4.517.8%1.1%1428
  71. #71DeepSeekDeepSeek V3.224.6%—1427
  72. #72DeepSeekDeepSeek V4 Flash38.6%16.6%1425
  73. #73OpenAIGPT-5.4 Nano28.3%9.3%1425
  74. #74MiniMaxMiniMax M2.729.6%0.6%1422
  75. #75AnthropicClaude Opus 412.3%0.3%1421
  76. #76xAIGrok 4.1 Fast (Reasoning)19.3%—1421
  77. #77xAIGrok 4.337.2%8.0%1419
  78. #78UpstageSolar Pro 429.2%—1412
  79. #79OpenAIGPT-5 Mini21.5%0.0%1405
  80. #80GoogleGemini 2.5 Flash12.1%1.1%1405
  81. #81AnthropicClaude Sonnet 410.7%0.3%1404
  82. #82AnthropicClaude Haiku 4.510.4%0.0%1400
  83. #83MiniMaxMiniMax M2.520.5%—1393
  84. #84Arcee AITrinity Large Thinking15.8%0.9%1384
  85. #85OpenAIGPT OSS 120B19.6%1.1%1381
  86. #86NVIDIANemotron 3 Super20.8%3.1%1373
  87. #87GoogleGemini 2.5 Flash Lite7.0%—1362
  88. #88NVIDIANemotron 3 Nano 30B A3B11.4%—1347
  89. #89OpenAIGPT-5 Nano9.5%—1346
  90. #90AmazonNova 2 Lite11.6%—1332
  91. #91IBMGranite 4.1 8B3.8%—1317
  92. #92MetaLlama 4 Maverick4.9%0.0%1316
  93. #93MetaLlama 4 Scout3.8%0.0%1309
  94. #94OpenAIGPT-4.14.2%—1303
Half of models ≤ 1452Measured by: User preference votesLast updated 2026-10-02

18 of the 34 AIs released in the last 60 days have no score here yet · Mistral Large 4, Perplexity Decider v1 27B, Clef Flash and 15 more

Source: Arena Intelligence Data sources & removal requests

What to keep in mind

  • It’s a preference vote, not a correctness check. A wrong but friendly answer can win.
  • Which questions got asked depends on who voted — they may differ from your use.
  • Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.