LiveBench Overall

In plain words · Overall score on questions refreshed monthly so they cannot be memorized

What does it measure?
An overall score from an exam that is freshly written every month. It spans every subject — reasoning, coding, agentic coding, maths, data analysis, language, instruction following — so it shows whether an AI is well-rounded rather than strong in one subject.
Who evaluates it, and how?
LiveBench, built by researchers at Abacus.AI, NYU and elsewhere, releases new questions every month from recent competitions, papers and news, reducing the chance that an AI has seen the questions during training (contamination). Only questions with fixed answers are used, graded by program rather than by another AI. This table takes the latest monthly results published at livebench.ai; an AI missing from the newest release is listed separately below the table with its last published score, unranked. A subject score is the average of its tasks. The overall score in this table is computed as the average of every task score we import (agentic coding included), so it can differ slightly from the six-subject average shown on the LiveBench site.
Reading the score
0–100. The questions are designed to be hard — even top AIs struggle to exceed 70 — and because they change monthly, avoid comparing scores from different dates.

Last updated: 2026-06-25

Rank

Top score83.4Claude Fable 5.1
Models tested50
Last updated2026.06.25
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 🥇AnthropicClaude Fable 5.153.4150183.4
  2. 🥈AnthropicClaude Fable 549.6150483.0
  3. 🥉OpenAIGPT-6 Astra52.7147782.2
  4. #4AnthropicClaude Opus 5.557.6150482.1
  5. #5MetaMuse Spark 1.348.1149481.6
  6. #6OpenAIGPT-6.1 Sol51.8148381.1
  7. #6DeepSeekDeepSeek V4.1 Flash39.5147481.1
  8. #8OpenAIGPT-5.6 Sol47.0148481.0
  9. #9OpenAIGPT-5.538.4147780.2
  10. #10AnthropicClaude Opus 550.8149080.1
  11. #11OpenAIGPT-6 Sol47.6145779.3
  12. #12Moonshot AIKimi K343.6148879.2
  13. #13GoogleGemini 3.7 Flash39.6148878.8
  14. #14xAIGrok 4.644.3145478.0
  15. #15OpenAIGPT-5.439.0147578.0
  16. #16MetaMuse Spark 1.239.6149478.0
  17. #17OpenAIGPT-5.6 Terra42.1146677.9
  18. #18AnthropicClaude Sonnet 5.556.0147177.8
  19. #19xAIGrok 4.746.4144377.4
  20. #20GoogleGemini 3.1 Pro29.7148777.0
  21. #21AnthropicClaude Opus 4.740.7149476.5
  22. #22AnthropicClaude Opus 4.841.8147576.2
  23. #23AlibabaQwen3.8 Flash39.8—76.2
  24. #24Z.aiGLM 5.344.8147976.1
  25. #25AnthropicClaude Sonnet 538.2146276.0
  26. #26GoogleGemini 3.8 Flash40.9149575.8
  27. #27xAIGrok 4.538.8146675.8
  28. #28MetaMuse Spark 1.133.7149175.3
  29. #29GoogleGemini 3.5 Flash33.6147674.6
  30. #30AnthropicClaude Opus 4.631.9149774.5
  31. #31GoogleGemini 3.6 Flash34.0148373.6
  32. #32OpenAIGPT-5.6 Luna37.3145373.6
  33. #33Z.aiGLM 5.233.7147673.2
  34. #34AlibabaQwen3.7 Max29.5147573.1
  35. #35AnthropicClaude Sonnet 4.630.1147273.0
  36. #36AnthropicClaude Opus 4.529.1147072.6
  37. #37OpenAIGPT-6 Luna38.1144372.0
  38. #38Z.aiGLM 5.3 Flash41.8147371.6
  39. #39DeepSeekDeepSeek V4 Pro36.0145871.6
  40. #40Moonshot AIKimi K2.627.0146170.5
  41. #41OpenAIGPT-5.4 Nano20.7140169.6
  42. #42AlibabaQwen3.6 Plus27.0144368.9
  43. #43Moonshot AIKimi K2.7 Code25.8—68.4
  44. #44xAIGrok Build 0.1——67.8
  45. #45NVIDIANemotron 3 Ultra22.9142667.4
  46. #46MiniMaxMiniMax M329.2144067.3
  47. #47OpenAIGPT-5.4 Mini24.1144766.4
  48. #48DeepSeekDeepSeek V4 Flash34.3143665.5
  49. #49GoogleGemini 3.5 Flash-Lite22.2145563.9
  50. #50xAIGrok 4.324.9144162.3
Half of models ≤ 75.9Last updated 2026-06-25

Scores from an earlier edition

These scores are not in the latest leaderboard (2026.06.25) — the model was dropped by the source or has not been refreshed yet. The last score we received is shown for reference, unranked.

  1. —
    OpenAI2025.08.07
    71.3
  2. —
    Z.ai2026.04.07
    70.6
  3. —
    Moonshot AI2026.01.27
    69.2
  4. —69.0
  5. —
    Z.ai2026.02.11
    68.7
  6. —
    MiniMax2026.03.18
    65.0
  7. —
    Google2026.04.02
    62.4
  8. —
    Google2026.03.03
    62.1
  9. —
    Anthropic2025.08.05
    61.4
  10. —
    OpenAI2025.08.07
    61.0
  11. —
    Anthropic2025.05.22
    60.6
  12. —
    Alibaba2026.04.27
    60.5
  13. —
    MiniMax2026.02.12
    60.3
  14. —60.1
  15. —
    Xiaomi2026.03.18
    58.4
  16. —
    Google2025.06.17
    57.5
  17. —
    Google2025.12.17
    54.4
  18. —
    Anthropic2025.09.29
    51.3
  19. —
    DeepSeek2025.12.01
    49.8
  20. —
    Z.ai2026.04.01
    48.8
  21. —
    OpenAI2025.08.07
    48.0
  22. —
    Google2025.06.17
    46.9
  23. —
    OpenAI2025.08.05
    46.4
  24. —
    Anthropic2025.10.15
    43.0
  25. —
    Google2025.09.25
    41.5
  26. —
    xAI2026.03.09
    37.9
  27. —
    NVIDIA2026.03.11
    32.0
  28. —
    xAI2025.11.19
    31.6
  29. —
    Arcee AI2026.04.01
    30.4

21 of the 37 AIs released in the last 60 days have no score here yet · Claude Haiku 5.5, Decider V1.1 27B, GPT-6 Luna Decisions and 18 more

Source: LiveBench Data sources & removal requests

What to keep in mind

  • This score comes from a fixed, predefined evaluation, so it may differ from what you get with your own question.
  • Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.