LiveBench Math

In plain words · How many freshly written math problems it solves

What does it measure?
Skill at competition-level maths problems and proofs — solving an unfamiliar problem through to the end.
Who evaluates it, and how?
LiveBench, built by researchers at Abacus.AI, NYU and elsewhere, releases new questions every month from recent competitions, papers and news, reducing the chance that an AI has seen the questions during training (contamination). Only questions with fixed answers are used, graded by program rather than by another AI. This table takes the latest monthly results published at livebench.ai; an AI missing from the newest release is listed separately below the table with its last published score, unranked. The maths subject consists of recent AMC/AIME problems (with reworded statements and reordered options), filling in missing steps of USAMO/IMO proofs, and harder synthetic calculation problems (AMPS).
Reading the score
0–100. The questions are designed to be hard — even top AIs struggle to exceed 70 — and because they change monthly, avoid comparing scores from different dates. Answers are fixed, so how friendly the explanation is does not affect the score.

Last updated: 2026-06-25

Rank

Top score97.0Claude Fable 5.1
Models tested50
Last updated2026.06.25
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 🥇AnthropicClaude Fable 5.197.0
  2. 🥈OpenAIGPT-6 Astra96.8
  3. 🥉AnthropicClaude Opus 5.596.8
  4. #4AnthropicClaude Sonnet 5.596.7
  5. #5OpenAIGPT-6.1 Sol96.5
  6. #6OpenAIGPT-6 Sol96.4
  7. #7OpenAIGPT-5.6 Sol96.2
  8. #8AnthropicClaude Fable 596.0
  9. #9MetaMuse Spark 1.396.0
  10. #10OpenAIGPT-5.595.9
  11. #11AnthropicClaude Opus 595.7
  12. #12xAIGrok 4.795.7
  13. #13OpenAIGPT-5.6 Terra94.9
  14. #14AnthropicClaude Opus 4.894.3
  15. #15OpenAIGPT-5.494.2
  16. #16GoogleGemini 3.7 Flash93.5
  17. #17DeepSeekDeepSeek V4.1 Flash93.3
  18. #18AnthropicClaude Sonnet 592.9
  19. #19AnthropicClaude Opus 4.792.8
  20. #20xAIGrok 4.692.6
  21. #21GoogleGemini 3.8 Flash91.6
  22. #22MetaMuse Spark 1.291.2
  23. #23GoogleGemini 3.1 Pro91.0
  24. #24OpenAIGPT-5.4 Nano91.0
  25. #25xAIGrok 4.590.8
  26. #26DeepSeekDeepSeek V4 Pro90.7
  27. #27AnthropicClaude Opus 4.590.4
  28. #28Z.aiGLM 5.289.8
  29. #29AnthropicClaude Opus 4.689.3
  30. #30OpenAIGPT-6 Luna89.1
  31. #31NVIDIANemotron 3 Ultra88.7
  32. #32GoogleGemini 3.5 Flash88.2
  33. #33Z.aiGLM 5.387.9
  34. #34OpenAIGPT-5.6 Luna87.2
  35. #35MetaMuse Spark 1.187.1
  36. #36AnthropicClaude Sonnet 4.687.0
  37. #37GoogleGemini 3.6 Flash86.4
  38. #38AlibabaQwen3.8 Flash85.8
  39. #39AlibabaQwen3.7 Max85.3
  40. #40Moonshot AIKimi K384.4
  41. #41xAIGrok 4.384.3
  42. #42Moonshot AIKimi K2.684.3
  43. #43AlibabaQwen3.6 Plus83.7
  44. #44Z.aiGLM 5.3 Flash81.2
  45. #45DeepSeekDeepSeek V4 Flash79.7
  46. #46Moonshot AIKimi K2.7 Code79.6
  47. #47OpenAIGPT-5.4 Mini78.5
  48. #48xAIGrok Build 0.178.4
  49. #49MiniMaxMiniMax M377.0
  50. #50GoogleGemini 3.5 Flash-Lite73.7
Half of models ≤ 90.8Last updated 2026-06-25

Scores from an earlier edition

These scores are not in the latest leaderboard (2026.06.25) — the model was dropped by the source or has not been refreshed yet. The last score we received is shown for reference, unranked.

  1. —
    Upstage
    Solar Pro 4Self-reported
    2026.07.22
    95.8
  2. —87.1
  3. —
    OpenAI2025.08.07
    86.2
  4. —
    Z.ai2026.04.07
    84.9
  5. —
    Moonshot AI2026.01.27
    84.9
  6. —83.7
  7. —
    Z.ai2026.02.11
    83.5
  8. —
    MiniMax2026.03.18
    80.5
  9. —
    Alibaba2026.04.27
    78.9
  10. —
    MiniMax2026.02.12
    77.4
  11. —
    Xiaomi2026.03.18
    77.0
  12. —
    OpenAI2025.08.07
    74.4
  13. —
    Google2026.04.02
    73.9
  14. —
    Google2026.03.03
    73.6
  15. —
    Anthropic2025.08.05
    73.2
  16. —
    Anthropic2025.05.22
    70.5
  17. —
    Z.ai2026.04.01
    70.4
  18. —
    OpenAI2025.08.05
    68.9
  19. —
    Google2025.06.17
    68.8
  20. —
    Google2025.06.17
    68.3
  21. —
    Google2025.12.17
    68.1
  22. —
    OpenAI2025.08.07
    64.7
  23. —
    DeepSeek2025.12.01
    64.0
  24. —
    Anthropic2025.09.29
    62.6
  25. —
    Google2025.09.25
    61.0
  26. —
    Anthropic2025.10.15
    58.0
  27. —
    xAI2026.03.09
    45.5
  28. —
    Arcee AI2026.04.01
    44.9
  29. —
    xAI2025.11.19
    38.9
  30. —
    NVIDIA2026.03.11
    36.4

21 of the 37 AIs released in the last 60 days have no score here yet · Claude Haiku 5.5, Decider V1.1 27B, GPT-6 Luna Decisions and 18 more

Source: LiveBench Data sources & removal requests

What to keep in mind

  • This score comes from a fixed, predefined evaluation, so it may differ from what you get with your own question.
  • Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.