FrontierCode

Widely citedHigher is better

Made by Cognition: for 150 hard issues from real open-source repositories, the agent must produce a fix good enough to merge. Grading covers behavior, not breaking existing features, build quality, and project conventions; the score is the average of five attempts on the 100 hardest tasks. Scores run from 0 to 100%. Higher is better.

Top score54.6%Claude Opus 5.5
Models tested36
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 🥇AnthropicClaude Opus 5.5Reasoning effort: Medium54.6%
  2. 🥈AnthropicClaude Fable 553.5%
  3. 🥉AnthropicClaude Opus 5Reasoning effort: Max53.4%
  4. #4OpenAIGPT-6 AstraReasoning effort: Max53.3%
  5. #5AnthropicClaude Sonnet 5.5Reasoning effort: Extra High52.1%
  6. #6AnthropicClaude Fable 5.1Reasoning effort: Medium50.9%
  7. #7OpenAIGPT-6.1 SolReasoning effort: Medium50.2%
  8. #8OpenAIGPT-6 SolReasoning effort: Max49.3%
  9. #9xAIGrok 4.648.0%
  10. #10xAIGrok 4.747.6%
  11. #11OpenAIGPT-5.6 Sol47.5%
  12. #12AnthropicClaude Opus 4.846.5%
  13. #13Moonshot AIKimi K344.2%
  14. #14GoogleGemini 3.7 Flash43.6%
  15. #15OpenAIGPT-5.543.0%
  16. #16AnthropicClaude Sonnet 542.7%
  17. #17xAIGrok 4.542.4%
  18. #18OpenAIGPT-6 LunaReasoning effort: Max42.4%
  19. #19OpenAIGPT-5.6 Terra41.3%
  20. #20GoogleGemini 3.8 FlashReasoning effort: Medium41.2%
  21. #21Z.aiGLM 5.3Reasoning effort: Max40.1%
  22. #22OpenAIGPT-5.6 Luna39.8%
  23. #23AnthropicClaude Opus 4.738.5%
  24. #24GoogleGemini 3.6 Flash34.4%
  25. #25Z.aiGLM 5.3 FlashReasoning effort: Max31.8%
  26. #26Moonshot AIKimi K2.7 Code30.1%
  27. #27DeepSeekDeepSeek V4 ProReasoning effort: High28.5%
  28. #28OpenAIGPT-5.4 Mini27.0%
  29. #29AnthropicClaude Opus 4.626.6%
  30. #30Z.aiGLM 5.2Reasoning effort: None24.5%
  31. #31AnthropicClaude Sonnet 4.624.3%
  32. #32DeepSeekDeepSeek V4 Flash18.8%
  33. #33MiniMaxMiniMax M314.7%
  34. #34NVIDIANemotron 3 Ultra13.6%
  35. #35AlibabaQwen3.7 Plus10.2%
  36. #36Mistral AIMistral Medium 3.58.0%
Half of models ≤ 41.9%Last updated 2026-10-07

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.