LiveBench Language

In plain words · How well it handles word puzzles and language fixes

What does it measure?
How precisely the AI handles language: word-grouping puzzles (Connections), spotting typos in paper abstracts, reordering shuffled movie plots — tasks needing reading and a feel for language.
Who evaluates it, and how?
LiveBench, built by researchers at Abacus.AI, NYU and elsewhere, releases new questions every month from recent competitions, papers and news, reducing the chance that an AI has seen the questions during training (contamination). Only questions with fixed answers are used, graded by program rather than by another AI. This table takes the latest monthly results published at livebench.ai; an AI missing from the newest release is listed separately below the table with its last published score, unranked. The language subject consists of Connections puzzles, typo removal in arXiv abstracts, and reordering plots of recent films.
Reading the score
0–100. The questions are designed to be hard — even top AIs struggle to exceed 70 — and because they change monthly, avoid comparing scores from different dates. It is English-based and not directly related to ability in other languages.

Last updated: 2026-06-25

Rank

Top score90.7Claude Fable 5
Models tested50
Last updated2026.06.25
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 🥇AnthropicClaude Fable 590.7
  2. 🥈AnthropicClaude Fable 5.189.5
  3. 🥉OpenAIGPT-6 Astra89.4
  4. #4AnthropicClaude Opus 588.7
  5. #5OpenAIGPT-6.1 Sol88.6
  6. #6GoogleGemini 3.8 Flash87.8
  7. #7OpenAIGPT-5.6 Sol87.7
  8. #8OpenAIGPT-5.587.4
  9. #9Moonshot AIKimi K385.5
  10. #10AnthropicClaude Opus 5.585.5
  11. #11GoogleGemini 3.7 Flash85.5
  12. #12GoogleGemini 3.1 Pro85.4
  13. #13OpenAIGPT-6 Sol85.3
  14. #14GoogleGemini 3.5 Flash84.6
  15. #15GoogleGemini 3.6 Flash83.9
  16. #16xAIGrok 4.683.7
  17. #17AnthropicClaude Sonnet 5.583.4
  18. #18AnthropicClaude Opus 4.683.3
  19. #19OpenAIGPT-5.6 Terra82.9
  20. #20xAIGrok 4.582.8
  21. #21MetaMuse Spark 1.382.8
  22. #22OpenAIGPT-5.482.6
  23. #23AnthropicClaude Opus 4.581.3
  24. #24DeepSeekDeepSeek V4.1 Flash81.2
  25. #25xAIGrok 4.780.1
  26. #26Z.aiGLM 5.379.9
  27. #27AlibabaQwen3.7 Max79.7
  28. #28AnthropicClaude Opus 4.879.7
  29. #29MetaMuse Spark 1.278.6
  30. #30DeepSeekDeepSeek V4 Pro78.1
  31. #31AnthropicClaude Opus 4.777.9
  32. #31Moonshot AIKimi K2.7 Code77.9
  33. #33Z.aiGLM 5.3 Flash77.3
  34. #34MiniMaxMiniMax M376.8
  35. #35Z.aiGLM 5.276.2
  36. #36AnthropicClaude Sonnet 4.676.1
  37. #37Moonshot AIKimi K2.675.1
  38. #38AlibabaQwen3.6 Plus75.0
  39. #39AnthropicClaude Sonnet 575.0
  40. #40AlibabaQwen3.8 Flash74.6
  41. #41MetaMuse Spark 1.174.3
  42. #42OpenAIGPT-6 Luna73.8
  43. #43xAIGrok 4.373.6
  44. #44OpenAIGPT-5.6 Luna72.6
  45. #45xAIGrok Build 0.172.5
  46. #46GoogleGemini 3.5 Flash-Lite71.8
  47. #47OpenAIGPT-5.4 Mini71.0
  48. #48NVIDIANemotron 3 Ultra70.8
  49. #49DeepSeekDeepSeek V4 Flash70.1
  50. #50OpenAIGPT-5.4 Nano62.5
Half of models ≤ 80.0Last updated 2026-06-25

Scores from an earlier edition

These scores are not in the latest leaderboard (2026.06.25) — the model was dropped by the source or has not been refreshed yet. The last score we received is shown for reference, unranked.

  1. —
    OpenAI2025.08.07
    80.7
  2. —
    Google2025.12.17
    78.7
  3. —77.7
  4. —
    Moonshot AI2026.01.27
    77.7
  5. —
    Z.ai2026.02.11
    77.5
  6. —
    Anthropic2025.09.29
    76.0
  7. —
    Google2025.06.17
    75.5
  8. —74.3
  9. —
    Google2026.03.03
    73.2
  10. —
    Anthropic2025.05.22
    72.9
  11. —
    Anthropic2025.08.05
    72.8
  12. —
    Z.ai2026.04.07
    71.8
  13. —
    Google2026.04.02
    71.3
  14. —
    OpenAI2025.08.07
    69.2
  15. —
    Xiaomi2026.03.18
    69.1
  16. —
    MiniMax2026.03.18
    66.8
  17. —
    DeepSeek2025.12.01
    64.2
  18. —
    Alibaba2026.04.27
    63.1
  19. —
    Z.ai2026.04.01
    62.3
  20. —
    Google2025.06.17
    62.3
  21. —
    Anthropic2025.10.15
    57.0
  22. —
    MiniMax2026.02.12
    55.1
  23. —
    Google2025.09.25
    52.0
  24. —
    xAI2025.11.19
    50.0
  25. —
    OpenAI2025.08.05
    48.6
  26. —
    OpenAI2025.08.07
    47.7
  27. —
    Arcee AI2026.04.01
    42.1
  28. —
    xAI2026.03.09
    42.0
  29. —
    NVIDIA2026.03.11
    30.0

21 of the 37 AIs released in the last 60 days have no score here yet · Claude Haiku 5.5, Decider V1.1 27B, GPT-6 Luna Decisions and 18 more

Source: LiveBench Data sources & removal requests

What to keep in mind

  • This score comes from a fixed, predefined evaluation, so it may differ from what you get with your own question.
  • Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.