SWE-Bench Pro

Widely citedHigher is better

A Scale AI test: the AI works on long software tasks from real open-source repositories and must pass hidden tests (731-task public set). Makers use different task sets and tools, so read each score's setting. Scores run from 0 to 100%. Higher is better.

This ranking is built only from scores the model makers published themselves. Test settings such as tool use and reasoning effort differ by maker, so check each score's setting before comparing models directly.

Top score89.9%Claude Opus 5.5
Models tested52
Last updated2026.09.22
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 🥇AnthropicClaude Opus 5.5Reasoning effort: Max89.9%
  2. 🥈AnthropicClaude Fable 5.1Reasoning effort: Max81.2%
  3. 🥉AnthropicClaude Fable 5Reasoning effort: Max80.0%
  4. #4AnthropicClaude Opus 5Reasoning effort: Max79.2%
  5. #5AnthropicClaude Opus 4.8Reasoning effort: Max69.2%
  6. #6TencentHy4 previewReasoning effort: Max65.7%
  7. #7xAIGrok 4.5Reasoning effort: High64.7%
  8. #8OpenAIGPT-5.6 Sol64.6%
  9. #9AnthropicClaude Opus 4.7Reasoning effort: Max64.3%
  10. #10OpenAIGPT-5.6 Terra63.4%
  11. #11AnthropicClaude Sonnet 5Reasoning effort: Max63.2%
  12. #12OpenAIGPT-5.6 Luna62.7%
  13. #13AlibabaQwen3.8 Flash62.5%
  14. #14Z.aiGLM 5.262.1%
  15. #15MetaMuse Spark 1.1Reasoning effort: Extra High61.5%
  16. #16AlibabaQwen3.7 Max60.6%
  17. #17MeituanLongCat 2.059.5%
  18. #18MiniMaxMiniMax M359.0%
  19. #19GoogleGemini 3.6 Flash58.7%
  20. #20OpenAIGPT-5.5Reasoning effort: Extra High58.6%
  21. #20Moonshot AIKimi K2.6Reasoning effort: Thinking58.6%
  22. #22Z.aiGLM-5.158.4%
  23. #23AnthropicClaude Sonnet 4.658.1%
  24. #24OpenAIGPT-5.4Reasoning effort: Extra High57.7%
  25. #25AlibabaQwen3.7 Plus57.6%
  26. #26AlibabaQwen3.6 Max57.3%
  27. #27XiaomiMiMo V2.5 Pro57.2%
  28. #28AlibabaQwen3.6 Plus56.6%
  29. #29StepfunStep 3.7 Flash56.3%
  30. #30MiniMaxMiniMax M2.756.2%
  31. #31XiaomiMiMo V2.556.1%
  32. #32MiniMaxMiniMax M2.555.4%
  33. #32DeepSeekDeepSeek V4 ProReasoning effort: Max55.4%
  34. #34GoogleGemini 3.5 Flash55.1%
  35. #34Z.aiGLM-555.1%
  36. #36XiaomiMiMo V2 Pro55.0%
  37. #36MetaMuse Spark55.0%
  38. #38OpenAIGPT-5.4 MiniReasoning effort: Extra High54.4%
  39. #39GoogleGemini 3.5 Flash-LiteReasoning effort: High54.2%
  40. #39GoogleGemini 3.1 ProReasoning effort: High54.2%
  41. #41AnthropicClaude Opus 4.653.4%
  42. #42DeepSeekDeepSeek V4 FlashReasoning effort: Max52.6%
  43. #43OpenAIGPT-5.4 NanoReasoning effort: Extra High52.4%
  44. #44AnthropicClaude Opus 4.5Reasoning effort: None52.0%
  45. #45Moonshot AIKimi K2.550.7%
  46. #46GoogleGemini 3 Flash49.6%
  47. #47AlibabaQwen3.6 Flash49.5%
  48. #47AlibabaQwen3.6 35B A3B49.5%
  49. #49PoolsideLaguna M.149.2%
  50. #50ByteDanceDola Seed 2.0 Pro46.9%
  51. #51CohereNorth Mini Code40.2%
  52. #52GoogleGemini 3.1 Flash Lite38.3%
Half of models ≤ 57.3%Measured by: Maker-reportedLast updated 2026-10-07

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.

Vendor-reported numbers use varying setups — treat them as indicative.