Agents' Last Exam
Widely citedHigher is better
Built by UC Berkeley with 250+ industry experts: an AI agent carries out long professional workflows from 55 computer-based occupations inside a virtual machine, using desktop apps, browsers, specialist software, and the command line. The score is the share of the 152 public tasks that receive full credit, with up to five hours per task. Scores run from 0 to 100%. Higher is better.
Other versions:Agents' Last Exam (CLI)
Top score39.5%Gemini 4 Argon
Models tested25
Last updated2026.09.30
Model release date (newest on the right)
Best score so farHigher on the chart is better
Can it handle customer service by the rules?τ³-Banking
Grok 4.6🥇
Grok 4.6🥇Can it run a small business alone for a year and make money?Vending-Bench 2
GPT-6 Astra🥇
GPT-6 Astra🥇
- 🥇GoogleGemini 4 ArgonReasoning effort: Max—39.5%$13,718
- 🥈AnthropicClaude Opus 5.5Reasoning effort: Max—38.2%$9,235
- 🥉OpenAIGPT-6 AstraReasoning effort: Max43.1%34.2%$15,515
- #4MetaMuse Spark 1.3Reasoning effort: Extra High50.5%32.2%—
- #4AnthropicClaude Opus 5Reasoning effort: High44.7%32.2%$11,182
- #4OpenAIGPT-6 SolReasoning effort: Extra High—32.2%$14,428
- #7OpenAIGPT-5.6 SolReasoning effort: Extra High44.3%30.6%$9,619
- #8OpenAIGPT-5.6 LunaReasoning effort: Extra High31.1%30.3%$4,095
- #9Moonshot AIKimi K3Reasoning effort: Max46.0%28.3%$5,165
- #10OpenAIGPT-5.6 TerraReasoning effort: Max40.2%28.0%$7,343
- #11xAIGrok 4.5Reasoning effort: High42.1%27.0%$3,887
- #11AnthropicClaude Opus 4.8Reasoning effort: Max34.2%27.0%$5,787
- #13OpenAIGPT-5.5Reasoning effort: Extra High39.0%26.6%$7,524
- #14OpenAIGPT-6 LunaReasoning effort: Max—25.0%—
- #15AnthropicClaude Opus 4.7Reasoning effort: High34.6%21.1%$10,937
- #16OpenAIGPT-5.439.6%20.5%$6,144
- #17Z.aiGLM 5.2Reasoning effort: Max34.6%20.4%$8,314
- #18GoogleGemini 3.1 Pro21.4%16.4%$3,774
- #19DeepSeekDeepSeek V4 Pro39.6%12.4%$3,285
- #20Z.aiGLM-5.113.6%11.5%$5,634
- #21Moonshot AIKimi K2.623.3%9.2%$6,205
- #22AlibabaQwen3.6 Plus20.8%8.6%$5,115
- #22XiaomiMiMo V2.58.7%8.6%—
- #24xAIGrok 4.312.4%7.2%$35
- #25MiniMaxMiniMax M2.79.9%5.9%—
Hard test · top score 39.5% · Half of models ≤ 26.6%Last updated 2026-10-07
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.
Vendor-reported numbers use varying setups — treat them as indicative.