Artificial Analysis Launches AutomationBench-AA Agent Leaderboard
- •Artificial Analysis launched AutomationBench-AA to test AI agent performance across 657 simulated SaaS workflow tasks.
- •Claude Fable 5 leads with a 48.6% score, followed by Claude Opus 4.8 at 48.5% and Gemini 3.5 Flash at 42.6%.
- •Finance workflows remain the most difficult to automate, with agents completing only one-third of objectives in that domain.
Artificial Analysis has launched AutomationBench-AA, an independent leaderboard testing the ability of AI agents to perform real-world SaaS workflow automation. Developed in partnership with Zapier, the benchmark evaluates model performance across 657 complex tasks spanning domains such as Finance, HR, Marketing, Operations, Sales, and Support. Agents operate within 40 simulated software environments—including Slack, Gmail, Google Sheets, Salesforce, Zendesk, Jira, and HubSpot—where they must navigate REST APIs, discover necessary endpoints, and process misleading information to achieve specific objectives. Models are scored against nearly 12,000 assertions designed to verify correct data handling while ensuring no business rules, known as guardrails, are violated.
Current performance leaders include Anthropic’s Claude Fable 5 and Opus 4.8, which achieved scores of 48.6% and 48.5% respectively. Google DeepMind's Gemini 3.5 Flash follows at 42.6%, and OpenAI's GPT-5.5 (xhigh) scores 42.1%. Performance metrics are adjusted by guardrail adherence; Gemini 3.5 Flash currently leads in efficiency, completing 15.0 objectives per guardrail violation, compared to 13.5 for Claude Opus 4.8 (max). While Gemini 3.5 Flash matches the performance of GPT-5.5 (xhigh) at a cost of $0.49 per task versus $1.32, it remains a competitive option. The leading open-weights model, GLM-5.2 (max) from Z.ai, sits at 27.8%, approximately 10 points behind the frontier models.
The evaluation revealed distinct operational strategies among leading agents. GPT-5.5 (xhigh) adopts an action-intensive approach, averaging 49 tool calls across 25 turns per task. Conversely, Claude Opus 4.8 (max) is more deliberate, requiring 35 tool calls over 14 turns with fewer guardrail violations at 0.55 per task. Domain-specific challenges also persist; Finance workflows remain the most difficult to automate, with agents completing roughly one-third of Finance objectives, compared to approximately 60% for Support and Operations tasks. Every model evaluated triggered at least some guardrail violations, with rates ranging from 0.46 per task for Gemini 3.5 Flash to 1.26 for Qwen3.7 Plus. Tasks are capped at 50 turns and graded deterministically based on final system states.