FrontierMath Tier 4 (2025-07)
Older versionHigher is better
The old version 1 (48 private problems) of Tier 4, the hardest research-level set in Epoch AI's FrontierMath. The AI can use Python and must give an exact answer. It was superseded by version 2 in June 2026, is no longer updated, and is not comparable with version 2 scores. Scores run from 0 to 100%. Higher is better.
A newer version is available: FrontierMath Tier 4 v2
Top score31.3%Claude Opus 4.8
Models tested29
Last updated2026.06.08
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelSettingReleasedPrice/1MScore
- 1max2026.05.27$2031.3%
- 2xhigh2026.03.05$1227.1%
- 3xhigh2026.04.16$2022.9%
- 4max2026.02.04$2022.9%
- 52026.02.19$9.5016.7%

- 62026.04.08—14.6%
- 7high2026.05.19$7.1314.6%

- 82026.04.20$2.8114.6%
- 9high2025.08.07$7.8112.5%
- 92026.04.07$3.6512.5%
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.