CharXiv (Reasoning)
Often citedHigher is better · Image-input models only
Questions about real charts taken from arXiv papers that require multi-step reasoning to answer (1,000-question validation set). Only models that accept images can take part. Scores run from 0 to 100%. Higher is better.
This ranking is built only from scores the model makers published themselves. Test settings such as tool use and reasoning effort differ by maker, so check each score's setting before comparing models directly.
Top score91.3%Kimi K3
Models tested31
Last updated2026.09.02
Model release date (newest on the right)
Best score so farHigher on the chart is better
Which AI gives the photo and chart answers people prefer most?Arena Vision
Claude Fable 5🥇
Claude Fable 5🥇Can it spot the wrong step from assembly photos?Furniture Assembly
Claude Opus 5.5🥇
Claude Opus 5.5🥇Can it draw a floor plan from a few photos of a home?Blueprint-Bench 2
Gemini 4 Argon🥇
Gemini 4 Argon🥇
- 🥇Moonshot AIKimi K3Reasoning effort: Max91.3%
- 🥈AnthropicClaude Opus 4.7Reasoning effort: Max91.0%
- 🥉AlibabaQwen3.8 Flash90.6%
- #4AnthropicClaude Opus 4.8Reasoning effort: Max89.9%
- #5GoogleGemini 3.6 Flash89.4%
- #5Z.aiGLM 5.3 FlashReasoning effort: Max89.4%
- #7MetaMuse Spark88.9%
- #8GoogleGemini 3.7 Flash88.7%
- #9MetaMuse Spark 1.1Reasoning effort: Extra High88.4%
- #10AnthropicClaude Sonnet 5Reasoning effort: Max88.3%
- #11Moonshot AIKimi K2.6Reasoning effort: Thinking86.7%
- #12GoogleGemini 3.8 Flash86.2%
- #13AlibabaQwen3.7 Plus85.9%
- #14GoogleGemini 3.5 Flash84.9%
- #15GoogleGemini 3.1 Pro83.3%
- #16AlibabaQwen3.6 Plus81.5%
- #17OpenAIGPT-5Reasoning effort: High81.1%
- #18XiaomiMiMo V2.581.0%
- #19ByteDanceDola Seed 2.0 Pro80.5%
- #20GoogleGemini 3 FlashReasoning effort: Thinking80.3%
- #21Moonshot AIKimi K2.578.7%
- #22AlibabaQwen3.6 Flash78.0%
- #22AlibabaQwen3.6 35B A3B78.0%
- #24AnthropicClaude Opus 4.6Reasoning effort: Max77.4%
- #24AnthropicClaude Sonnet 4.6Reasoning effort: Max77.4%
- #26GoogleGemini 3.5 Flash-LiteReasoning effort: High76.5%
- #27GoogleGemini 3.1 Flash Lite75.6%
- #28OpenAIGPT-5 MiniReasoning effort: High75.5%
- #29GoogleGemini 2.5 ProReasoning effort: Thinking69.6%
- #30AnthropicClaude Opus 4.5Reasoning effort: Max68.7%
- #31BaiduERNIE 5.0 Thinking67.1%
Top 6 near 91.3% · hard to separate the best · Half of models ≤ 81.5%Measured by: Maker-reportedLast updated 2026-10-07
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.
Vendor-reported numbers use varying setups — treat them as indicative.