GSO-Bench
Higher is better
Given a real codebase and a performance test, the agent must speed the code up as much as an expert developer's optimization did (102 tasks across 10 codebases in five languages). The score is the share of tasks where one attempt reaches at least 95% of the expert's speed-up while staying correct. Scores run from 0 to 100%. Higher is better.
Top score88.2%Claude Fable 5.1
Models tested18
Last updated2026.09.27
Model release date (newest on the right)
Best score so farHigher on the chart is better
Can it operate a computer itself to finish hours-long development jobs?Terminal-Bench 4.0
Claude Sonnet 5.5🥇
Claude Sonnet 5.5🥇Which AI gives the coding answers people prefer most?Arena Coding
Gemini 4 Argon🥇
Gemini 4 Argon🥇Can it run and fix code until the task is done?LiveBench Agentic Coding
DeepSeek V4.1 Flash🥇
DeepSeek V4.1 Flash🥇
- 🥇AnthropicClaude Fable 5.188.2%
- 🥈OpenAIGPT-6 Astra79.4%
- 🥉AnthropicClaude Fable 578.4%
- #4OpenAIGPT-5.6 Sol76.5%
- #5AnthropicClaude Opus 4.847.1%
- #6AnthropicClaude Opus 4.7Reasoning effort: High44.1%
- #7AnthropicClaude Opus 4.6Reasoning effort: High41.2%
- #8OpenAIGPT-5.5Reasoning effort: Extra High40.2%
- #9AnthropicClaude Sonnet 537.3%
- #10OpenAIGPT-5.4Reasoning effort: Extra High31.4%
- #11AnthropicClaude Opus 4.526.5%
- #12GoogleGemini 3.1 Pro22.6%
- #13AnthropicClaude Sonnet 4.514.7%
- #14GoogleGemini 3 Flash9.8%
- #15AnthropicClaude Opus 46.9%
- #16OpenAIGPT-5Reasoning effort: High6.9%
- #17AnthropicClaude Sonnet 44.9%
- #18GoogleGemini 2.5 Pro3.9%
Half of models ≤ 34.3%Last updated 2026-10-07
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.