FrontierMath: Erdős Problems

Higher is better

68 conjectures from problems posed or studied by Paul Erdős, all still unsolved as of August 2026: the AI must prove or disprove each one with a machine-checked formal proof (Lean). It gets one attempt per problem, with limits of $300 and 72 hours. Scores run from 0 to 100%. Higher is better.

Top score2.9%GPT-6 Astra
Models tested5
Last updated2026.09.01
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    OpenAI2026.09.04 · max
    2.9%
  2. 2
    Anthropic2026.09.01 · max
    0.0%
  3. 2
    OpenAI2026.04.23 · xhigh
    0.0%
  4. 2
    Anthropic2026.06.09 · max
    0.0%
  5. 2
    OpenAI2026.07.09 · max
    0.0%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.