CursorBench
Often citedHigher is better
Tasks collected from real sessions of the Cursor coding tool: ambiguous, multi-file work such as understanding code, finding bugs, editing, refactoring, and code review. Cursor runs it on its own product's tasks, and scores from different versions are not comparable. Scores run from 0 to 100%. Higher is better.
Top score57.8%Claude Opus 5.5
Models tested14
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelSettingReleasedPrice/1MScore
- 1Max2026.09.22$1657.8%
- 2Max2026.09.28$855.5%
- 3Max2026.09.01$4051.8%
- 4Max2026.07.24$2046.6%
- 5Extra High2026.09.21$546.3%
- 6Max2026.08.18$3.6042.6%
- 7Max2026.07.09$1641.7%
- 8Max2026.09.02$3.5041.6%
- 9Extra High2026.08.12$541.4%
- 10Max2026.07.09$9.5041.3%
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.