LiveBench Math
In plain words · How many freshly written math problems it solves
- What does it measure?
- Skill at competition-level maths problems and proofs — solving an unfamiliar problem through to the end.
- Who evaluates it, and how?
- LiveBench, built by researchers at Abacus.AI, NYU and elsewhere, releases new questions every month from recent competitions, papers and news, reducing the chance that an AI has seen the questions during training (contamination). Only questions with fixed answers are used, graded by program rather than by another AI. This table takes the latest monthly results published at livebench.ai; an AI missing from the newest release is listed separately below the table with its last published score, unranked. The maths subject consists of recent AMC/AIME problems (with reworded statements and reordered options), filling in missing steps of USAMO/IMO proofs, and harder synthetic calculation problems (AMPS).
- Reading the score
- 0–100. The questions are designed to be hard — even top AIs struggle to exceed 70 — and because they change monthly, avoid comparing scores from different dates. Answers are fixed, so how friendly the explanation is does not affect the score.
Last updated: 2026-06-25
Rank
Top score97.0Claude Fable 5.1
Models tested50
Last updated2026.06.25
Model release date (newest on the right)
Best score so farHigher on the chart is better
Can it solve problems that stump experts across fields?Humanity's Last Exam
Claude Opus 5.5🥇
Claude Opus 5.5🥇Can it solve research-level physics problems?CritPt
GPT-5.6 Sol🥇
GPT-5.6 Sol🥇Which AI gives the math answers people prefer most?Arena Math
Gemini 4 Argon🥇
Gemini 4 Argon🥇
- 🥇AnthropicClaude Fable 5.197.0
- 🥈OpenAIGPT-6 Astra96.8
- 🥉AnthropicClaude Opus 5.596.8
- #4AnthropicClaude Sonnet 5.596.7
- #5OpenAIGPT-6.1 Sol96.5
- #6OpenAIGPT-6 Sol96.4
- #7OpenAIGPT-5.6 Sol96.2
- #8AnthropicClaude Fable 596.0
- #9MetaMuse Spark 1.396.0
- #10OpenAIGPT-5.595.9
- #11AnthropicClaude Opus 595.7
- #12xAIGrok 4.795.7
- #13OpenAIGPT-5.6 Terra94.9
- #14AnthropicClaude Opus 4.894.3
- #15OpenAIGPT-5.494.2
- #16GoogleGemini 3.7 Flash93.5
- #17DeepSeekDeepSeek V4.1 Flash93.3
- #18AnthropicClaude Sonnet 592.9
- #19AnthropicClaude Opus 4.792.8
- #20xAIGrok 4.692.6
- #21GoogleGemini 3.8 Flash91.6
- #22MetaMuse Spark 1.291.2
- #23GoogleGemini 3.1 Pro91.0
- #24OpenAIGPT-5.4 Nano91.0
- #25xAIGrok 4.590.8
- #26DeepSeekDeepSeek V4 Pro90.7
- #27AnthropicClaude Opus 4.590.4
- #28Z.aiGLM 5.289.8
- #29AnthropicClaude Opus 4.689.3
- #30OpenAIGPT-6 Luna89.1
- #31NVIDIANemotron 3 Ultra88.7
- #32GoogleGemini 3.5 Flash88.2
- #33Z.aiGLM 5.387.9
- #34OpenAIGPT-5.6 Luna87.2
- #35MetaMuse Spark 1.187.1
- #36AnthropicClaude Sonnet 4.687.0
- #37GoogleGemini 3.6 Flash86.4
- #38AlibabaQwen3.8 Flash85.8
- #39AlibabaQwen3.7 Max85.3
- #40Moonshot AIKimi K384.4
- #41xAIGrok 4.384.3
- #42Moonshot AIKimi K2.684.3
- #43AlibabaQwen3.6 Plus83.7
- #44Z.aiGLM 5.3 Flash81.2
- #45DeepSeekDeepSeek V4 Flash79.7
- #46Moonshot AIKimi K2.7 Code79.6
- #47OpenAIGPT-5.4 Mini78.5
- #48xAIGrok Build 0.178.4
- #49MiniMaxMiniMax M377.0
- #50GoogleGemini 3.5 Flash-Lite73.7
Half of models ≤ 90.8Last updated 2026-06-25
Scores from an earlier edition
These scores are not in the latest leaderboard (2026.06.25) — the model was dropped by the source or has not been refreshed yet. The last score we received is shown for reference, unranked.
- —2026.07.22$0.9795.8
- —2026.03.31$2.1987.1
- —2025.08.07$7.8186.2
- —2026.04.07$3.6584.9
- —2026.01.27$2.2784.9
- —2025.09.15$0.4283.7
- —2026.02.11$2.4083.5
- —2026.03.18$0.9780.5
- —2026.04.27$0.8978.9
- —2026.02.12$0.9777.4
- —2026.03.18$2.5077.0

- —2025.08.07$1.5674.4
- —2026.04.02$0.3473.9

- —2026.03.03$1.1973.6

- —2025.08.05$6073.2
- —2025.05.22$1270.5
- —2026.04.01$3.3070.4
- —2025.08.05$0.4068.9
- —2025.06.17$1.9568.8

- —2025.06.17$7.8168.3

- —2025.12.17$2.3868.1

- —2025.08.07$0.3164.7
- —2025.12.01$0.8364.0
- —2025.09.29$1262.6
- —2025.09.25$0.3361.0

- —2025.10.15$458.0
- —2026.03.09$2.1945.5
- —2026.04.01$0.6644.9
- —2025.11.19$0.4238.9
- —2026.03.11$0.3436.4
21 of the 37 AIs released in the last 60 days have no score here yet · Claude Haiku 5.5, Decider V1.1 27B, GPT-6 Luna Decisions and 18 more
Source: LiveBench Data sources & removal requestsWhat to keep in mind
- This score comes from a fixed, predefined evaluation, so it may differ from what you get with your own question.
- Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.