Furniture Assembly
Higher is better · Image-input models only
Made by Epoch AI: from 60 photos of partly or fully assembled IKEA furniture plus the official assembly manual, the model judges whether the build is correct so far and, if not, which step went wrong and how. Only models that accept images can take part. Scores run from 0 to 100%. Higher is better.
Top score83.3%Claude Opus 5.5
Models tested27
Last updated2026.09.29
Model release date (newest on the right)
Best score so farHigher on the chart is better
Which AI gives the photo and chart answers people prefer most?Arena Vision
Claude Fable 5🥇
Claude Fable 5🥇Can it draw a floor plan from a few photos of a home?Blueprint-Bench 2
Gemini 4 Argon🥇
Gemini 4 Argon🥇
- 🥇AnthropicClaude Opus 5.5Reasoning effort: Max—83.3%51.2%
- 🥈OpenAIGPT-6.1 SolReasoning effort: Max128580.0%—
- 🥈OpenAIGPT-6 AstraReasoning effort: Max128080.0%49.7%
- #4AnthropicClaude Sonnet 5.5Reasoning effort: Max129075.0%—
- #5AnthropicClaude Fable 5.1Reasoning effort: Max132070.0%41.9%
- #6AnthropicClaude Opus 5Reasoning effort: Max—60.8%30.4%
- #7OpenAIGPT-6 SolReasoning effort: Max—58.3%36.9%
- #8OpenAIGPT-5.6 SolReasoning effort: Max127956.7%33.6%
- #9OpenAIGPT-5.6 TerraReasoning effort: Max127154.2%30.8%
- #10OpenAIGPT-5.5Reasoning effort: Extra High129744.2%36.2%
- #10OpenAIGPT-6 LunaReasoning effort: Max—44.2%31.2%
- #12AnthropicClaude Opus 4.8Reasoning effort: Max128942.5%14.5%
- #12OpenAIGPT-5.6 LunaReasoning effort: Max126042.5%22.6%
- #14xAIGrok 4.6Reasoning effort: Extra High126540.0%33.2%
- #15OpenAIGPT-5.4Reasoning effort: Extra High130237.5%27.1%
- #16AnthropicClaude Fable 5Reasoning effort: Max132535.8%38.6%
- #17Moonshot AIKimi K3Reasoning effort: Max—34.2%29.5%
- #18AnthropicClaude Opus 4.7Reasoning effort: Max131633.3%24.5%
- #19GoogleGemini 3.8 FlashReasoning effort: High131631.7%38.6%
- #20AnthropicClaude Opus 4.6Reasoning effort: Max131328.3%—
- #20AnthropicClaude Opus 4.5Reasoning effort: Budget 64K—28.3%—
- #22GoogleGemini 3.7 FlashReasoning effort: High131526.7%—
- #22GoogleGemini 3.1 ProReasoning effort: High129626.7%26.5%
- #24GoogleGemini 3.6 FlashReasoning effort: High129723.3%31.2%
- #25xAIGrok 4.5Reasoning effort: High128822.5%27.3%
- #26Moonshot AIKimi K2.6128221.7%3.9%
- #27xAIGrok 4.7Reasoning effort: Extra High—20.8%32.5%
Half of models ≤ 40.0%Last updated 2026-10-07
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.