Agents' Last Exam (CLI)

Often citedHigher is better

The Linux-only part of Agents' Last Exam, 105 tasks that run in Linux virtual machines, meant for comparing command-line-only agents. Its score is not the same number as the full 152-task score. Scores run from 0 to 100%. Higher is better.

Other versions:Agents' Last Exam
Top score34.3%Claude Opus 5.5
Models tested25
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    Anthropic2026.09.22 · claude_code · thinking-max
    34.3%
  2. 2
    Meta2026.09.02 · codex · reasoning-xhigh
    33.3%
  3. 2
    OpenAI2026.09.04 · codex · reasoning-max
    33.3%
  4. 4
    OpenAI2026.09.22 · codex · reasoning-high
    32.4%
  5. 5
    OpenAI2026.07.09 · codex · reasoning-xhigh
    29.5%
  6. 5
    Anthropic2026.07.24 · claude_code · thinking-high
    29.5%
  7. 7
    OpenAI2026.07.09 · codex · reasoning-xhigh
    28.6%
  8. 8
    Moonshot AI2026.07.17 · claude_code · thinking-max
    27.6%
  9. 9
    OpenAI2026.04.23 · codex · reasoning-xhigh
    27.1%
  10. 10
    OpenAI2026.07.09 · codex · reasoning-xhigh
    26.2%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.