ProofBench (Vals)
Higher is better
A formal-proof test by Vals AI: for 100 theorems at graduate or advanced-undergraduate level (analysis, algebra, probability, number theory, and more), the model must write a Lean 4 proof that passes the machine checker, with no partial credit. It is not the same as Google DeepMind's similarly named test. Scores run from 0 to 100%. Higher is better.
Top score100.0%Claude Fable 5.1
Models tested65
Model release date (newest on the right)
Best score so farHigher on the chart is better
Can it solve problems that stump experts across fields?Humanity's Last Exam
Claude Opus 5.5🥇
Claude Opus 5.5🥇Can it solve research-level physics problems?CritPt
GPT-5.6 Sol🥇
GPT-5.6 Sol🥇Which AI gives the math answers people prefer most?Arena Math
Gemini 4 Argon🥇
Gemini 4 Argon🥇
- 🥇AnthropicClaude Fable 5.1Reasoning effort: Max100.0%
- 🥇AnthropicClaude Opus 5.5Reasoning effort: Max100.0%
- 🥇AnthropicClaude Sonnet 5.5Reasoning effort: Max100.0%
- #4GoogleGemini 4 Argon99.0%
- #4OpenAIGPT-6.1 Sol99.0%
- #4AnthropicClaude Opus 5Reasoning effort: Max99.0%
- #4OpenAIGPT-6 Astra99.0%
- #8AnthropicClaude Fable 5Reasoning effort: Max95.0%
- #9Moonshot AIKimi K387.0%
- #10OpenAIGPT-6 Sol83.0%
- #10OpenAIGPT-5.6 SolReasoning effort: Max83.0%
- #12AnthropicClaude Sonnet 5Reasoning effort: Max77.0%
- #13TencentHy4 preview75.0%
- #14OpenAIGPT-5.6 TerraReasoning effort: Extra High74.0%
- #15XiaomiMiMo-V2.6-Pro70.0%
- #16AnthropicClaude Opus 4.8Reasoning effort: Max69.0%
- #17OpenAIGPT-6 Luna64.0%
- #18XiaomiMiMo-V2.6-Flash63.0%
- #19OpenAIGPT-5.6 LunaReasoning effort: Max60.0%
- #20MetaMuse Spark 1.3Reasoning effort: Max58.0%
- #20GoogleGemini 3.7 Flash58.0%
- #22DeepSeekDeepSeek V4 Flash56.0%
- #22OpenAIGPT-5.4Reasoning effort: Extra High56.0%
- #24DeepSeekDeepSeek V4.1 Flash54.0%
- #24AnthropicClaude Opus 4.7Reasoning effort: Max54.0%
- #26xAIGrok 4.651.0%
- #27AnthropicClaude Opus 4.6Reasoning effort: Max50.0%
- #27OpenAIGPT-5.5Reasoning effort: Extra High50.0%
- #27DeepSeekDeepSeek V4 Pro50.0%
- #30Z.aiGLM 5.3Reasoning effort: Max49.0%
- #31GoogleGemini 3.8 Flash48.0%
- #32AnthropicClaude Sonnet 4.6Reasoning effort: Max45.0%
- #33MetaMuse Spark 1.243.0%
- #34MetaMuse Spark 1.139.0%
- #35AnthropicClaude Opus 4.536.0%
- #35GoogleGemini 3.6 Flash36.0%
- #37Z.aiGLM 5.2Reasoning effort: Max35.0%
- #38xAIGrok 4.734.0%
- #39GoogleGemini 3.5 FlashReasoning effort: High31.0%
- #39xAIGrok 4.5Reasoning effort: High31.0%
- #41GoogleGemini 3.1 Pro26.0%
- #41AlibabaQwen3.7 Max26.0%
- #43Z.aiGLM-5.122.2%
- #44XiaomiMiMo V2.5 Pro22.0%
- #45Z.aiGLM 5.3 FlashReasoning effort: Max21.0%
- #45OpenAIGPT-5.4 MiniReasoning effort: Extra High21.0%
- #47AnthropicClaude Sonnet 4.519.0%
- #48OpenAIGPT-5Reasoning effort: High18.0%
- #48MiniMaxMiniMax M318.0%
- #50MetaMuse Spark17.0%
- #51XiaomiMiMo V2.516.0%
- #51Moonshot AIKimi K2.616.0%
- #53GoogleGemini 3 Flash15.0%
- #54xAIGrok 4.20 (Reasoning)14.0%
- #55GoogleGemini 3.5 Flash-Lite13.0%
- #56OpenAIGPT-5 NanoReasoning effort: High12.0%
- #57xAIGrok 4.3Reasoning effort: High11.0%
- #58OpenAIGPT-5 MiniReasoning effort: High9.0%
- #58Mistral AIMistral Medium 3.59.0%
- #60OpenAIGPT-5.4 NanoReasoning effort: High5.0%
- #61MiniMaxMiniMax M2.54.0%
- #61xAIGrok 4.1 Fast (Reasoning)4.0%
- #63MiniMaxMiniMax M2.73.0%
- #64NVIDIANemotron 3 Ultra2.0%
- #65PoolsideLaguna M.10.0%
Top 7 near 100.0% · hard to separate the best · Half of models ≤ 43.0%Last updated 2026-10-07
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.