FrontierMath Tier 4 (2025-07)

Older versionHigher is better

The old version 1 (48 private problems) of Tier 4, the hardest research-level set in Epoch AI's FrontierMath. The AI can use Python and must give an exact answer. It was superseded by version 2 in June 2026, is no longer updated, and is not comparable with version 2 scores. Scores run from 0 to 100%. Higher is better.

A newer version is available: FrontierMath Tier 4 v2

Top score31.3%Claude Opus 4.8
Models tested29
Last updated2026.06.08
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    Anthropic2026.05.27 · max
    31.3%
  2. 2
    OpenAI2026.03.05 · xhigh
    27.1%
  3. 3
    Anthropic2026.04.16 · xhigh
    22.9%
  4. 4
    Anthropic2026.02.04 · max
    22.9%
  5. 5
    Google2026.02.19
    16.7%
  6. 6
    Meta2026.04.08
    14.6%
  7. 7
    Google2026.05.19 · high
    14.6%
  8. 8
    Moonshot AI2026.04.20
    14.6%
  9. 9
    OpenAI2025.08.07 · high
    12.5%
  10. 9
    Z.ai2026.04.07
    12.5%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.