Arena Mathematical
In plain words · Which AI people preferred on expert-level mathematics questions
- What does it measure?
- Which AI's answers were preferred on expert-level questions in Mathematical & Statistics. Examples: checking proofs, choosing statistical models, optimisation problems. These are questions the people who actually do that work would ask, not general ones.
- Who evaluates it, and how?
- On Arena Intelligence, a user enters a question, two anonymous AIs answer, and the user votes for the better one. Over 82 million votes are aggregated with a statistical model (Bradley-Terry) into an Elo-style score, and this table takes the 'style-controlled' version from the official dataset, which statistically removes the bias toward long, nicely formatted answers. An AI classifier judges each prompt's reasoning depth and expertise and tags only about 5.5% as 'expert' prompts, which are then divided into occupational categories based on the US Bureau of Labor Statistics classification (SOC), launched November 2025. This table counts only prompts classified as 'Mathematical & Statistics'.
- Reading the score
- This is a head-to-head record, not an exam score (accuracy %). It is computed from the probability of being preferred over other AIs, so there is no maximum, and when a new AI enters, existing scores shift slightly. Typical values are 1,000–1,500; a 100-point gap means the leader is preferred about 64% of the time, a 200-point gap about 76%. Correctness is not counted — a wrong but friendly answer can win — so if accuracy matters, also look at accuracy tests (GPQA, AIME, etc.). Gaps of a few dozen points are effectively the same tier even if the rank differs. This category has few prompts, so top ranks change often; look at score bands rather than ranks.
Last updated: 2026-10-02
Rank
Top score1542Gemini 4 Argon
Models tested94
Last updated2026.10.02
Model release date (newest on the right)
Best score so farHigher on the chart is better
Can it solve problems that stump experts across fields?Humanity's Last Exam
Claude Opus 5.5🥇
Claude Opus 5.5🥇Can it solve research-level physics problems?CritPt
GPT-5.6 Sol🥇
GPT-5.6 Sol🥇Which AI gives the math answers people prefer most?Arena Math
Gemini 4 Argon🥇
Gemini 4 Argon🥇
- 🥇GoogleGemini 4 Argon1542
- 🥈AnthropicClaude Fable 5.11523
- 🥉AnthropicClaude Fable 51522
- #4AnthropicClaude Opus 5.51521
- #5AnthropicClaude Opus 51518
- #6AnthropicClaude Opus 4.61513
- #7OpenAIGPT-5.6 Sol1513
- #8GoogleGemini 3.8 Flash1512
- #9Moonshot AIKimi K31509
- #10DeepSeekDeepSeek V4.1 Flash1503
- #11XiaomiMiMo-V2.6-Pro1501
- #12MetaMuse Spark 1.11500
- #13Z.aiGLM 5.31499
- #14OpenAIGPT-5.41499
- #15OpenAIGPT-5.51499
- #16AnthropicClaude Opus 4.71499
- #17OpenAIGPT-6 Astra1498
- #18GoogleGemini 3.7 Flash1495
- #19Z.aiGLM 5.3 Flash1495
- #20MetaMuse Spark 1.21494
- #21Z.aiGLM 5.21494
- #22xAIGrok 4.51493
- #23MetaMuse Spark 1.31493
- #24XiaomiMiMo V2.5 Pro1491
- #25AnthropicClaude Opus 4.81490
- #26OpenAIGPT-5.6 Terra1489
- #27GoogleGemini 3.5 Flash1488
- #28GoogleGemini 3.6 Flash1487
- #29AlibabaQwen3.6 Max1487
- #30AnthropicClaude Sonnet 51487
- #31OpenAIGPT-5.6 Luna1487
- #32xAIGrok 4.61485
- #33Moonshot AIKimi K2.61485
- #34AnthropicClaude Sonnet 4.61485
- #35GoogleGemini 3.1 Pro1484
- #36Moonshot AIKimi K2.51477
- #37AnthropicClaude Opus 4.51475
- #38Z.aiGLM-5.11474
- #39AlibabaQwen3.7 Plus1474
- #40GoogleGemma 4 31B1473
- #41XiaomiMiMo V2 Pro1471
- #41MetaMuse Spark1471
- #43GoogleGemini 3 Flash1468
- #44OpenAIGPT-6 Luna1468
- #45AlibabaQwen3.7 Max1466
- #46xAIGrok 4.20 (Reasoning)1460
- #47DeepSeekDeepSeek V4 Pro1460
- #48Z.aiGLM-51458
- #49XiaomiMiMo-V2.6-Flash1457
- #50AlibabaQwen3.6 Plus1455
- #51xAIGrok 4.201455
- #52AlibabaQwen3.5 397B A17B1454
- #53XiaomiMiMo V2.51454
- #54OpenAIGPT-6 Sol1454
- #55AnthropicClaude Sonnet 4.51452
- #55OpenAIGPT-5.4 Mini1452
- #55TencentHy31452
- #58NVIDIANemotron 3 Ultra1451
- #59AnthropicClaude Opus 4.11451
- #60xAIGrok 4.71450
- #61ByteDanceDola Seed 2.0 Pro1448
- #62GoogleGemini 3.5 Flash-Lite1448
- #62MiniMaxMiniMax M31448
- #64GoogleGemini 2.5 Pro1447
- #65Z.aiGLM 5V Turbo1446
- #66OpenAIGPT-51442
- #67MiniMaxMiniMax M2.71441
- #68MeituanLongcat Flash Chat1440
- #69Mistral AIMistral Medium 3.51439
- #70BaiduERNIE 5.0 Thinking1435
- #71DeepSeekDeepSeek V4 Flash1434
- #72AnthropicClaude Haiku 4.51433
- #73DeepSeekDeepSeek V3.21432
- #74GoogleGemini 3.1 Flash Lite1432
- #75xAIGrok 4.31431
- #76OpenAIGPT-5.4 Nano1431
- #77AnthropicClaude Opus 41426
- #78xAIGrok 4.1 Fast (Reasoning)1423
- #79UpstageSolar Pro 41421
- #80GoogleGemini 2.5 Flash1416
- #81AnthropicClaude Sonnet 41413
- #82OpenAIGPT-5 Mini1407
- #83MiniMaxMiniMax M2.51399
- #83NVIDIANemotron 3 Super1399
- #85OpenAIGPT OSS 120B1384
- #86Arcee AITrinity Large Thinking1376
- #87GoogleGemini 2.5 Flash Lite1370
- #88OpenAIGPT-5 Nano1355
- #89NVIDIANemotron 3 Nano 30B A3B1354
- #90AmazonNova 2 Lite1347
- #91IBMGranite 4.1 8B1326
- #92MetaLlama 4 Maverick1319
- #93MetaLlama 4 Scout1314
- #94OpenAIGPT-4.11309
Half of models ≤ 1459Measured by: User preference votesLast updated 2026-10-02
21 of the 37 AIs released in the last 60 days have no score here yet · Claude Haiku 5.5, Decider V1.1 27B, GPT-6 Luna Decisions and 18 more
Source: Arena Intelligence Data sources & removal requestsWhat to keep in mind
- It’s a preference vote, not a correctness check. A wrong but friendly answer can win.
- Which questions got asked depends on who voted — they may differ from your use.
- Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.