CursorBench

Often citedHigher is better

Tasks collected from real sessions of the Cursor coding tool: ambiguous, multi-file work such as understanding code, finding bugs, editing, refactoring, and code review. Cursor runs it on its own product's tasks, and scores from different versions are not comparable. Scores run from 0 to 100%. Higher is better.

Top score57.8%Claude Opus 5.5
Models tested14
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    Anthropic2026.09.22 · Max
    57.8%
  2. 2
    Anthropic2026.09.28 · Max
    55.5%
  3. 3
    Anthropic2026.09.01 · Max
    51.8%
  4. 4
    Anthropic2026.07.24 · Max
    46.6%
  5. 5
    xAI2026.09.21 · Extra High
    46.3%
  6. 6
    Z.ai2026.08.18 · Max
    42.6%
  7. 7
    OpenAI2026.07.09 · Max
    41.7%
  8. 8
    Meta2026.09.02 · Max
    41.6%
  9. 9
    xAI2026.08.12 · Extra High
    41.4%
  10. 10
    OpenAI2026.07.09 · Max
    41.3%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.