Agents' Last Exam

Widely citedHigher is better

Built by UC Berkeley with 250+ industry experts: an AI agent carries out long professional workflows from 55 computer-based occupations inside a virtual machine, using desktop apps, browsers, specialist software, and the command line. The score is the share of the 152 public tasks that receive full credit, with up to five hours per task. Scores run from 0 to 100%. Higher is better.

Top score39.5%Gemini 4 Argon
Models tested25
Last updated2026.09.30
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 🥇GoogleGemini 4 ArgonReasoning effort: Max—39.5%$13,718
  2. 🥈AnthropicClaude Opus 5.5Reasoning effort: Max—38.2%$9,235
  3. 🥉OpenAIGPT-6 AstraReasoning effort: Max43.1%34.2%$15,515
  4. #4MetaMuse Spark 1.3Reasoning effort: Extra High50.5%32.2%—
  5. #4AnthropicClaude Opus 5Reasoning effort: High44.7%32.2%$11,182
  6. #4OpenAIGPT-6 SolReasoning effort: Extra High—32.2%$14,428
  7. #7OpenAIGPT-5.6 SolReasoning effort: Extra High44.3%30.6%$9,619
  8. #8OpenAIGPT-5.6 LunaReasoning effort: Extra High31.1%30.3%$4,095
  9. #9Moonshot AIKimi K3Reasoning effort: Max46.0%28.3%$5,165
  10. #10OpenAIGPT-5.6 TerraReasoning effort: Max40.2%28.0%$7,343
  11. #11xAIGrok 4.5Reasoning effort: High42.1%27.0%$3,887
  12. #11AnthropicClaude Opus 4.8Reasoning effort: Max34.2%27.0%$5,787
  13. #13OpenAIGPT-5.5Reasoning effort: Extra High39.0%26.6%$7,524
  14. #14OpenAIGPT-6 LunaReasoning effort: Max—25.0%—
  15. #15AnthropicClaude Opus 4.7Reasoning effort: High34.6%21.1%$10,937
  16. #16OpenAIGPT-5.439.6%20.5%$6,144
  17. #17Z.aiGLM 5.2Reasoning effort: Max34.6%20.4%$8,314
  18. #18GoogleGemini 3.1 Pro21.4%16.4%$3,774
  19. #19DeepSeekDeepSeek V4 Pro39.6%12.4%$3,285
  20. #20Z.aiGLM-5.113.6%11.5%$5,634
  21. #21Moonshot AIKimi K2.623.3%9.2%$6,205
  22. #22AlibabaQwen3.6 Plus20.8%8.6%$5,115
  23. #22XiaomiMiMo V2.58.7%8.6%—
  24. #24xAIGrok 4.312.4%7.2%$35
  25. #25MiniMaxMiniMax M2.79.9%5.9%—
Hard test · top score 39.5% · Half of models ≤ 26.6%Last updated 2026-10-07

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.

Vendor-reported numbers use varying setups — treat them as indicative.