Agents' Last Exam (CLI)
Often citedHigher is better
The Linux-only part of Agents' Last Exam, 105 tasks that run in Linux virtual machines, meant for comparing command-line-only agents. Its score is not the same number as the full 152-task score. Scores run from 0 to 100%. Higher is better.
Other versions:Agents' Last Exam
Top score34.3%Claude Opus 5.5
Models tested25
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelSettingReleasedPrice/1MScore
- 1claude_code · thinking-max2026.09.22$1634.3%
- 2codex · reasoning-xhigh2026.09.02$3.5033.3%
- 2codex · reasoning-max2026.09.04$4033.3%
- 4codex · reasoning-high2026.09.22$832.4%
- 5codex · reasoning-xhigh2026.07.09$0.9529.5%
- 5claude_code · thinking-high2026.07.24$2029.5%
- 7codex · reasoning-xhigh2026.07.09$1628.6%
- 8claude_code · thinking-max2026.07.17$1127.6%
- 9codex · reasoning-xhigh2026.04.23$2427.1%
- 10codex · reasoning-xhigh2026.07.09$9.5026.2%
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.