LiveBench Reasoning
In plain words · How many freshly written logic puzzles it solves each month
- What does it measure?
- Skill at logic puzzles: drawing conclusions from several people's true/false statements, placement puzzles with interlocking constraints, spatial reasoning about shapes — thinking things through rather than knowing facts.
- Who evaluates it, and how?
- LiveBench, built by researchers at Abacus.AI, NYU and elsewhere, releases new questions every month from recent competitions, papers and news, reducing the chance that an AI has seen the questions during training (contamination). Only questions with fixed answers are used, graded by program rather than by another AI. This table takes the latest monthly results published at livebench.ai; an AI missing from the newest release is listed separately below the table with its last published score, unranked. The reasoning subject consists of truth/lie deduction (Web of Lies), Zebra puzzles and spatial reasoning tasks.
- Reading the score
- 0–100. The questions are designed to be hard — even top AIs struggle to exceed 70 — and because they change monthly, avoid comparing scores from different dates. It measures pure logic, not common sense or knowledge.
Last updated: 2026-06-25
Rank
Top score92.7GPT-6 Astra
Models tested50
Last updated2026.06.25
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelReleasedPrice/1MScore
- 12026.09.04$4092.7
- 22026.09.01$4091.7
- 32026.07.09$1691.7
- 42026.09.29$891.6
- 52026.07.24$2091.2
- 62026.07.17$1190.7
- 72026.09.22$1690.7
- 82026.07.09$9.5090.6
- 92026.08.12$590.5
- 102026.08.05$3.5090.0
21 of the 37 AIs released in the last 60 days have no score here yet · Claude Haiku 5.5, Decider V1.1 27B, GPT-6 Luna Decisions and 18 more
Source: LiveBench Data sources & removal requestsWhat to keep in mind
- This score comes from a fixed, predefined evaluation, so it may differ from what you get with your own question.
- Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.