MCP-Atlas

Often citedHigher is better

A Scale AI test of whether the AI can complete 1,000 tasks using 36 real tool servers connected through MCP. Scores run from 0 to 100%. Higher is better.

This ranking is built only from scores the model makers published themselves. Test settings such as tool use and reasoning effort differ by maker, so check each score's setting before comparing models directly.

Top score90.3%Muse Spark 1.2
Models tested37
Last updated2026.09.20
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 🥇MetaMuse Spark 1.2Reasoning effort: Extra High90.3%
  2. 🥈MetaMuse Spark 1.1Reasoning effort: Extra High88.1%
  3. 🥉AnthropicClaude Opus 585.8%
  4. #4StepfunStep-5Reasoning effort: High85.6%
  5. #5Moonshot AIKimi K3Reasoning effort: Max84.2%
  6. #6TencentHy4 previewReasoning effort: Max83.7%
  7. #7GoogleGemini 3.5 Flash83.6%
  8. #8AnthropicClaude Fable 583.3%
  9. #9AnthropicClaude Opus 4.8Reasoning effort: Max82.2%
  10. #9MetaMuse Spark82.2%
  11. #11AnthropicClaude Opus 4.7Reasoning effort: High79.7%
  12. #12GoogleGemini 3.1 Pro78.2%
  13. #13Z.aiGLM 5.276.8%
  14. #14AlibabaQwen3.7 Max76.4%
  15. #15Moonshot AIKimi K2.7 CodeReasoning effort: Thinking76.0%
  16. #16AnthropicClaude Opus 4.675.8%
  17. #17OpenAIGPT-5.5Reasoning effort: Extra High75.3%
  18. #18MiniMaxMiniMax M374.2%
  19. #18DeepSeekDeepSeek V4 ProReasoning effort: High74.2%
  20. #20AlibabaQwen3.6 Plus74.1%
  21. #21AlibabaQwen3.7 Plus73.2%
  22. #22Z.aiGLM-5.171.8%
  23. #23OpenAIGPT-5.4Reasoning effort: Extra High70.6%
  24. #24Moonshot AIKimi K2.6Reasoning effort: Thinking69.4%
  25. #25DeepSeekDeepSeek V4 FlashReasoning effort: Max69.0%
  26. #26Z.aiGLM-5Reasoning effort: Thinking67.8%
  27. #27AlibabaQwen3.6 Flash62.8%
  28. #27AlibabaQwen3.6 35B A3B62.8%
  29. #29AnthropicClaude Opus 4.5Reasoning effort: None62.3%
  30. #30GoogleGemini 3 Flash62.0%
  31. #31UpstageSolar Pro 461.4%
  32. #32AnthropicClaude Sonnet 4.6Reasoning effort: Max61.3%
  33. #33OpenAIGPT-5.4 MiniReasoning effort: Extra High57.7%
  34. #34OpenAIGPT-5.4 NanoReasoning effort: Extra High56.1%
  35. #35MiniMaxMiniMax M2.749.4%
  36. #36AnthropicClaude Sonnet 4.543.8%
  37. #37GoogleGemini 2.5 ProReasoning effort: Thinking8.8%
Half of models ≤ 74.2%Measured by: Maker-reportedLast updated 2026-10-07

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.

Vendor-reported numbers use varying setups — treat them as indicative.