IMO-AnswerBench

Often citedHigher is better

400 International Mathematical Olympiad–level problems from Google DeepMind, answered with a short final answer. There is no public leaderboard, so only maker-reported values exist. Scores run from 0 to 100%. Higher is better.

This ranking is built only from scores the model makers published themselves. Test settings such as tool use and reasoning effort differ by maker, so check each score's setting before comparing models directly.

Top score92.3%Nemotron 3 Ultra
Models tested20
Last updated2026.08.07
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 🥇NVIDIANemotron 3 Ultra92.3%
  2. 🥈Z.aiGLM 5.291.0%
  3. 🥉AlibabaQwen3.7 MaxReasoning effort: Extra High90.0%
  4. #4DeepSeekDeepSeek V4 ProReasoning effort: Max89.8%
  5. #5ByteDanceDola Seed 2.0 Pro89.3%
  6. #6DeepSeekDeepSeek V4 FlashReasoning effort: Max88.4%
  7. #7AlibabaQwen3.7 Plus86.0%
  8. #7Moonshot AIKimi K2.6Reasoning effort: Thinking86.0%
  9. #9TencentHy384.3%
  10. #10AlibabaQwen3.6 Plus83.8%
  11. #10Z.aiGLM-5.183.8%
  12. #12Motif TechnologiesMotif 383.2%
  13. #13Z.aiGLM-582.5%
  14. #14Moonshot AIKimi K2.5Reasoning effort: Thinking81.8%
  15. #14MeituanLongCat 2.081.8%
  16. #16AlibabaQwen3.6 Flash78.9%
  17. #16AlibabaQwen3.6 35B A3B78.9%
  18. #18LG AI ResearchK-EXAONE 2.0Reasoning effort: Thinking78.6%
  19. #19DeepSeekDeepSeek V3.2Reasoning effort: Thinking78.3%
  20. #20LG AI ResearchK-EXAONEReasoning effort: Thinking76.3%
Half of models ≤ 83.8%Measured by: Maker-reportedLast updated 2026-10-07

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.

Vendor-reported numbers use varying setups — treat them as indicative.