ProofBench (Vals)

Higher is better

A formal-proof test by Vals AI: for 100 theorems at graduate or advanced-undergraduate level (analysis, algebra, probability, number theory, and more), the model must write a Lean 4 proof that passes the machine checker, with no partial credit. It is not the same as Google DeepMind's similarly named test. Scores run from 0 to 100%. Higher is better.

Top score100.0%Claude Fable 5.1
Models tested65
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 🥇AnthropicClaude Fable 5.1Reasoning effort: Max100.0%
  2. 🥇AnthropicClaude Opus 5.5Reasoning effort: Max100.0%
  3. 🥇AnthropicClaude Sonnet 5.5Reasoning effort: Max100.0%
  4. #4GoogleGemini 4 Argon99.0%
  5. #4OpenAIGPT-6.1 Sol99.0%
  6. #4AnthropicClaude Opus 5Reasoning effort: Max99.0%
  7. #4OpenAIGPT-6 Astra99.0%
  8. #8AnthropicClaude Fable 5Reasoning effort: Max95.0%
  9. #9Moonshot AIKimi K387.0%
  10. #10OpenAIGPT-6 Sol83.0%
  11. #10OpenAIGPT-5.6 SolReasoning effort: Max83.0%
  12. #12AnthropicClaude Sonnet 5Reasoning effort: Max77.0%
  13. #13TencentHy4 preview75.0%
  14. #14OpenAIGPT-5.6 TerraReasoning effort: Extra High74.0%
  15. #15XiaomiMiMo-V2.6-Pro70.0%
  16. #16AnthropicClaude Opus 4.8Reasoning effort: Max69.0%
  17. #17OpenAIGPT-6 Luna64.0%
  18. #18XiaomiMiMo-V2.6-Flash63.0%
  19. #19OpenAIGPT-5.6 LunaReasoning effort: Max60.0%
  20. #20MetaMuse Spark 1.3Reasoning effort: Max58.0%
  21. #20GoogleGemini 3.7 Flash58.0%
  22. #22DeepSeekDeepSeek V4 Flash56.0%
  23. #22OpenAIGPT-5.4Reasoning effort: Extra High56.0%
  24. #24DeepSeekDeepSeek V4.1 Flash54.0%
  25. #24AnthropicClaude Opus 4.7Reasoning effort: Max54.0%
  26. #26xAIGrok 4.651.0%
  27. #27AnthropicClaude Opus 4.6Reasoning effort: Max50.0%
  28. #27OpenAIGPT-5.5Reasoning effort: Extra High50.0%
  29. #27DeepSeekDeepSeek V4 Pro50.0%
  30. #30Z.aiGLM 5.3Reasoning effort: Max49.0%
  31. #31GoogleGemini 3.8 Flash48.0%
  32. #32AnthropicClaude Sonnet 4.6Reasoning effort: Max45.0%
  33. #33MetaMuse Spark 1.243.0%
  34. #34MetaMuse Spark 1.139.0%
  35. #35AnthropicClaude Opus 4.536.0%
  36. #35GoogleGemini 3.6 Flash36.0%
  37. #37Z.aiGLM 5.2Reasoning effort: Max35.0%
  38. #38xAIGrok 4.734.0%
  39. #39GoogleGemini 3.5 FlashReasoning effort: High31.0%
  40. #39xAIGrok 4.5Reasoning effort: High31.0%
  41. #41GoogleGemini 3.1 Pro26.0%
  42. #41AlibabaQwen3.7 Max26.0%
  43. #43Z.aiGLM-5.122.2%
  44. #44XiaomiMiMo V2.5 Pro22.0%
  45. #45Z.aiGLM 5.3 FlashReasoning effort: Max21.0%
  46. #45OpenAIGPT-5.4 MiniReasoning effort: Extra High21.0%
  47. #47AnthropicClaude Sonnet 4.519.0%
  48. #48OpenAIGPT-5Reasoning effort: High18.0%
  49. #48MiniMaxMiniMax M318.0%
  50. #50MetaMuse Spark17.0%
  51. #51XiaomiMiMo V2.516.0%
  52. #51Moonshot AIKimi K2.616.0%
  53. #53GoogleGemini 3 Flash15.0%
  54. #54xAIGrok 4.20 (Reasoning)14.0%
  55. #55GoogleGemini 3.5 Flash-Lite13.0%
  56. #56OpenAIGPT-5 NanoReasoning effort: High12.0%
  57. #57xAIGrok 4.3Reasoning effort: High11.0%
  58. #58OpenAIGPT-5 MiniReasoning effort: High9.0%
  59. #58Mistral AIMistral Medium 3.59.0%
  60. #60OpenAIGPT-5.4 NanoReasoning effort: High5.0%
  61. #61MiniMaxMiniMax M2.54.0%
  62. #61xAIGrok 4.1 Fast (Reasoning)4.0%
  63. #63MiniMaxMiniMax M2.73.0%
  64. #64NVIDIANemotron 3 Ultra2.0%
  65. #65PoolsideLaguna M.10.0%
Top 7 near 100.0% · hard to separate the best · Half of models ≤ 43.0%Last updated 2026-10-07

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.