Arena Math
In plain words · Which AI people preferred for math questions
- What does it measure?
- Preference on maths questions (calculation, worked solutions, proofs). Because people choose rather than checking against an answer key, solutions that are easy to read have an advantage.
- Who evaluates it, and how?
- On Arena Intelligence, a user enters a question, two anonymous AIs answer, and the user votes for the better one. Over 82 million votes are aggregated with a statistical model (Bradley-Terry) into an Elo-style score, and this table takes the 'style-controlled' version from the official dataset, which statistically removes the bias toward long, nicely formatted answers. An AI classifier selects prompts that actively apply mathematical concepts — calculations and derivations.
- Reading the score
- This is a head-to-head record, not an exam score (accuracy %). It is computed from the probability of being preferred over other AIs, so there is no maximum, and when a new AI enters, existing scores shift slightly. Typical values are 1,000–1,500; a 100-point gap means the leader is preferred about 64% of the time, a 200-point gap about 76%. Correctness is not counted — a wrong but friendly answer can win — so if accuracy matters, also look at accuracy tests (GPQA, AIME, etc.). Gaps of a few dozen points are effectively the same tier even if the rank differs. It measures something different from accuracy tests (AIME, MATH-500): not whether the answer was right, but whether people liked it.
Last updated: 2026-10-02
Rank
Top score1530Gemini 4 Argon
Models tested94
Last updated2026.10.02
Model release date (newest on the right)
Best score so farHigher on the chart is better
Can it solve problems that stump experts across fields?Humanity's Last Exam
Claude Opus 5.5🥇
Claude Opus 5.5🥇Can it solve research-level physics problems?CritPt
GPT-5.6 Sol🥇
GPT-5.6 Sol🥇
- 🥇GoogleGemini 4 Argon57.1%27.1%1530
- 🥈AnthropicClaude Fable 555.5%28.6%1522
- 🥉AnthropicClaude Fable 5.159.1%31.1%1518
- #4GoogleGemini 3.8 Flash47.8%18.3%1518
- #5AnthropicClaude Opus 554.9%29.1%1516
- #6AnthropicClaude Opus 5.561.4%31.7%1511
- #7MetaMuse Spark 1.348.7%26.0%1509
- #8AnthropicClaude Opus 4.639.9%—1507
- #9GoogleGemini 3.7 Flash47.9%14.3%1503
- #10Z.aiGLM 5.3 Flash39.9%15.4%1503
- #11DeepSeekDeepSeek V4.1 Flash39.2%14.3%1503
- #12GoogleGemini 3.6 Flash40.8%10.6%1501
- #13Moonshot AIKimi K346.9%23.4%1501
- #14OpenAIGPT-5.545.8%27.1%1500
- #15Z.aiGLM 5.342.3%19.1%1494
- #16OpenAIGPT-5.6 Sol49.5%32.3%1494
- #17MetaMuse Spark 1.146.2%15.1%1492
- #18OpenAIGPT-5.443.7%23.4%1492
- #19AnthropicClaude Opus 4.742.3%12.0%1492
- #20OpenAIGPT-6 Astra54.7%31.7%1490
- #21GoogleGemini 3.1 Pro47.0%17.7%1488
- #21Z.aiGLM 5.241.1%20.9%1488
- #23XiaomiMiMo-V2.6-Pro49.4%26.6%1487
- #24AlibabaQwen3.7 Max40.5%13.4%1484
- #25GoogleGemini 3.5 Flash42.7%13.1%1482
- #26XiaomiMiMo V2.5 Pro35.7%4.0%1481
- #27Moonshot AIKimi K2.637.5%8.0%1480
- #28OpenAIGPT-5.6 Terra42.9%30.0%1479
- #29OpenAIGPT-5.6 Luna39.5%20.6%1477
- #30AnthropicClaude Opus 4.848.7%20.9%1477
- #31AlibabaQwen3.6 Max30.8%—1476
- #32Z.aiGLM-5.130.1%4.6%1475
- #33GoogleGemini 3 Flash36.6%—1474
- #34AnthropicClaude Sonnet 541.3%16.9%1474
- #35xAIGrok 4.542.7%15.4%1473
- #36XiaomiMiMo-V2.6-Flash35.1%12.0%1472
- #37Moonshot AIKimi K2.530.7%3.1%1471
- #38GoogleGemma 4 31B23.6%1.4%1469
- #39MetaMuse Spark 1.245.5%17.7%1468
- #40MetaMuse Spark40.7%11.3%1466
- #41AlibabaQwen3.7 Plus35.6%9.1%1466
- #42xAIGrok 4.20 (Reasoning)34.5%—1465
- #43AnthropicClaude Sonnet 4.633.6%3.1%1464
- #44AnthropicClaude Opus 4.530.1%—1463
- #45OpenAIGPT-6 Luna38.5%19.4%1460
- #46AlibabaQwen3.6 Plus27.8%2.9%1454
- #47AlibabaQwen3.5 397B A17B29.0%—1452
- #48OpenAIGPT-6 Sol47.9%30.9%1451
- #49XiaomiMiMo V2 Pro30.4%—1451
- #50ByteDanceDola Seed 2.0 Pro——1450
- #51xAIGrok 4.2027.9%—1449
- #52xAIGrok 4.644.1%19.7%1448
- #53DeepSeekDeepSeek V4 Pro41.0%18.0%1445
- #54xAIGrok 4.743.1%18.0%1444
- #55Z.aiGLM-529.3%—1444
- #56AnthropicClaude Opus 4.112.5%—1443
- #57GoogleGemini 3.5 Flash-Lite18.8%0.0%1442
- #58NVIDIANemotron 3 Ultra28.4%3.1%1442
- #59GoogleGemini 2.5 Pro22.5%2.0%1439
- #60OpenAIGPT-5.4 Mini28.1%10.0%1439
- #61XiaomiMiMo V2.527.2%3.7%1438
- #62Z.aiGLM 5V Turbo17.1%—1437
- #63GoogleGemini 3.1 Flash Lite17.2%1.1%1436
- #64MeituanLongcat Flash Chat5.8%—1435
- #65BaiduERNIE 5.0 Thinking13.3%—1435
- #66OpenAIGPT-528.5%12.6%1434
- #67MiniMaxMiniMax M339.0%3.7%1432
- #68TencentHy333.5%—1431
- #69Mistral AIMistral Medium 3.513.8%0.0%1429
- #70AnthropicClaude Sonnet 4.517.8%1.1%1428
- #71DeepSeekDeepSeek V3.224.6%—1427
- #72DeepSeekDeepSeek V4 Flash38.6%16.6%1425
- #73OpenAIGPT-5.4 Nano28.3%9.3%1425
- #74MiniMaxMiniMax M2.729.6%0.6%1422
- #75AnthropicClaude Opus 412.3%0.3%1421
- #76xAIGrok 4.1 Fast (Reasoning)19.3%—1421
- #77xAIGrok 4.337.2%8.0%1419
- #78UpstageSolar Pro 429.2%—1412
- #79OpenAIGPT-5 Mini21.5%0.0%1405
- #80GoogleGemini 2.5 Flash12.1%1.1%1405
- #81AnthropicClaude Sonnet 410.7%0.3%1404
- #82AnthropicClaude Haiku 4.510.4%0.0%1400
- #83MiniMaxMiniMax M2.520.5%—1393
- #84Arcee AITrinity Large Thinking15.8%0.9%1384
- #85OpenAIGPT OSS 120B19.6%1.1%1381
- #86NVIDIANemotron 3 Super20.8%3.1%1373
- #87GoogleGemini 2.5 Flash Lite7.0%—1362
- #88NVIDIANemotron 3 Nano 30B A3B11.4%—1347
- #89OpenAIGPT-5 Nano9.5%—1346
- #90AmazonNova 2 Lite11.6%—1332
- #91IBMGranite 4.1 8B3.8%—1317
- #92MetaLlama 4 Maverick4.9%0.0%1316
- #93MetaLlama 4 Scout3.8%0.0%1309
- #94OpenAIGPT-4.14.2%—1303
Half of models ≤ 1452Measured by: User preference votesLast updated 2026-10-02
18 of the 34 AIs released in the last 60 days have no score here yet · Mistral Large 4, Perplexity Decider v1 27B, Clef Flash and 15 more
Source: Arena Intelligence Data sources & removal requestsWhat to keep in mind
- It’s a preference vote, not a correctness check. A wrong but friendly answer can win.
- Which questions got asked depends on who voted — they may differ from your use.
- Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.