Humanity's Last Exam (with tools)

Often citedHigher is better

Humanity's Last Exam (2,500 expert-written questions across more than 100 subjects) taken with tools such as search and code execution. It is a different setting from the no-tools score. Scores run from 0 to 100%. Higher is better.

This ranking is built only from scores the model makers published themselves. Test settings such as tool use and reasoning effort differ by maker, so check each score's setting before comparing models directly.

Top score67.7%Claude Opus 5.5
Models tested44
Last updated2026.09.22
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 🥇AnthropicClaude Opus 5.5Reasoning effort: Max67.7%
  2. 🥈AnthropicClaude Fable 5.1Reasoning effort: Max65.0%
  3. 🥉AnthropicClaude Opus 5Reasoning effort: Max64.7%
  4. #4DeepSeekDeepSeek V4.1 FlashReasoning effort: Max63.9%
  5. #4AnthropicClaude Fable 563.9%
  6. #6Z.aiGLM 5.3Reasoning effort: Max62.5%
  7. #7MetaMuse Spark 1.1Reasoning effort: Extra High62.1%
  8. #8StepfunStep-5Reasoning effort: High59.4%
  9. #9OpenAIGPT-5.4 ProReasoning effort: Extra High58.7%
  10. #10AnthropicClaude Opus 4.8Reasoning effort: Max57.9%
  11. #11AnthropicClaude Sonnet 5Reasoning effort: Max57.4%
  12. #12OpenAIGPT-6 Astra57.2%
  13. #13Moonshot AIKimi K3Reasoning effort: Max56.0%
  14. #14Moonshot AIKimi K2.6Reasoning effort: Thinking55.5%
  15. #15TencentHy4 previewReasoning effort: Max55.4%
  16. #16Z.aiGLM 5.3 FlashReasoning effort: Max55.3%
  17. #17Z.aiGLM 5.254.7%
  18. #17AnthropicClaude Opus 4.7Reasoning effort: Max54.7%
  19. #19ByteDanceDola Seed 2.0 Pro54.2%
  20. #20AlibabaQwen3.7 Max53.5%
  21. #21AnthropicClaude Opus 4.6Reasoning effort: Max53.0%
  22. #22Z.aiGLM-5.152.3%
  23. #23OpenAIGPT-5.5Reasoning effort: Extra High52.2%
  24. #24OpenAIGPT-5.4Reasoning effort: Extra High52.1%
  25. #25Moonshot AIKimi K2.5Reasoning effort: Thinking51.8%
  26. #26GoogleGemini 3.1 ProReasoning effort: High51.4%
  27. #27AlibabaQwen3.6 Plus50.6%
  28. #28MetaMuse SparkReasoning effort: Thinking50.4%
  29. #28Z.aiGLM-550.4%
  30. #30StepfunStep 3.7 Flash49.7%
  31. #31AnthropicClaude Sonnet 4.6Reasoning effort: Max49.0%
  32. #32DeepSeekDeepSeek V4 ProReasoning effort: Max48.2%
  33. #33XiaomiMiMo V2.5 Pro48.0%
  34. #34DeepSeekDeepSeek V4 FlashReasoning effort: Max45.1%
  35. #35GoogleGemini 3 FlashReasoning effort: Thinking43.5%
  36. #36AnthropicClaude Opus 4.5Reasoning effort: None43.2%
  37. #37OpenAIGPT-5.4 MiniReasoning effort: Extra High41.5%
  38. #38DeepSeekDeepSeek V3.2Reasoning effort: Thinking40.8%
  39. #39OpenAIGPT-5.4 NanoReasoning effort: Extra High37.7%
  40. #40NVIDIANemotron 3 Ultra37.4%
  41. #41AnthropicClaude Sonnet 4.533.6%
  42. #42GoogleGemma 4 31B26.5%
  43. #43NVIDIANemotron 3 Super22.8%
  44. #44OpenAIGPT OSS 120BReasoning effort: High19.0%
Half of models ≤ 52.3%Measured by: Maker-reportedLast updated 2026-10-07

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.

Vendor-reported numbers use varying setups — treat them as indicative.