SimpleBench

Often citedHigher is better

213 six-option questions that are easy for ordinary people but trip up AI: spatial and temporal reasoning, social awareness, and linguistic trick questions. Each question is run five times and averaged; non-specialist human participants scored 83.7%. Scores run from 0 to 100%. Higher is better.

Top score81.9%Claude Fable 5
Models tested40
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    Anthropic2026.06.09 · max
    81.9%
  2. 2
    Anthropic2026.07.24
    80.6%
  3. 3
    Google2026.02.19
    79.6%
  4. 4
    Google2026.05.19
    76.7%
  5. 5
    xAI2026.08.12
    75.9%
  6. 6
    Meta2026.08.05
    74.5%
  7. 7
    OpenAI2026.03.05
    74.1%
  8. 8
    OpenAI2026.07.09 · xhigh
    71.7%
  9. 9
    Alibaba2026.05.21
    70.4%
  10. 10
    xAI2026.07.08
    70.0%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.