SimpleBench
Often citedHigher is better
213 six-option questions that are easy for ordinary people but trip up AI: spatial and temporal reasoning, social awareness, and linguistic trick questions. Each question is run five times and averaged; non-specialist human participants scored 83.7%. Scores run from 0 to 100%. Higher is better.
Top score81.9%Claude Fable 5
Models tested40
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelSettingReleasedPrice/1MScore
- 1max2026.06.09$4081.9%
- 22026.07.24$2080.6%
- 32026.02.19$9.5079.6%

- 42026.05.19$7.1376.7%

- 52026.08.12$575.9%
- 62026.08.05$3.5074.5%
- 72026.03.05$14374.1%
- 8xhigh2026.07.09$1671.7%
- 92026.05.21$3.6970.4%
- 102026.07.08$570.0%
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.