GSO-Bench

Higher is better

Given a real codebase and a performance test, the agent must speed the code up as much as an expert developer's optimization did (102 tasks across 10 codebases in five languages). The score is the share of tasks where one attempt reaches at least 95% of the expert's speed-up while staying correct. Scores run from 0 to 100%. Higher is better.

Top score88.2%Claude Fable 5.1
Models tested18
Last updated2026.09.27
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 🥇AnthropicClaude Fable 5.188.2%
  2. 🥈OpenAIGPT-6 Astra79.4%
  3. 🥉AnthropicClaude Fable 578.4%
  4. #4OpenAIGPT-5.6 Sol76.5%
  5. #5AnthropicClaude Opus 4.847.1%
  6. #6AnthropicClaude Opus 4.7Reasoning effort: High44.1%
  7. #7AnthropicClaude Opus 4.6Reasoning effort: High41.2%
  8. #8OpenAIGPT-5.5Reasoning effort: Extra High40.2%
  9. #9AnthropicClaude Sonnet 537.3%
  10. #10OpenAIGPT-5.4Reasoning effort: Extra High31.4%
  11. #11AnthropicClaude Opus 4.526.5%
  12. #12GoogleGemini 3.1 Pro22.6%
  13. #13AnthropicClaude Sonnet 4.514.7%
  14. #14GoogleGemini 3 Flash9.8%
  15. #15AnthropicClaude Opus 46.9%
  16. #16OpenAIGPT-5Reasoning effort: High6.9%
  17. #17AnthropicClaude Sonnet 44.9%
  18. #18GoogleGemini 2.5 Pro3.9%
Half of models ≤ 34.3%Last updated 2026-10-07

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.