FrontierMath Tiers 1-3 (2025-02)
Older versionHigher is better
The old version 1 (290 private problems) of FrontierMath Tiers 1-3, original math problems written by mathematicians, from advanced undergraduate to early research level. It was replaced in June 2026 by version 2, which fixed errors, and is not directly comparable with version 2 scores. Scores run from 0 to 100%. Higher is better.
A newer version is available: FrontierMath Tiers 1-3 v2
Top score50.0%GPT-5.4 Pro
Models tested32
Last updated2026.06.08
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelSettingReleasedPrice/1MScore
- 1xhigh2026.03.05$14350.0%
- 2xhigh2026.03.05$1247.6%
- 3max2026.05.27$2047.2%
- 4xhigh2026.04.16$2043.8%
- 5max2026.02.04$2040.7%
- 62026.04.08—39.0%
- 7high2026.05.19$7.1339.0%

- 72026.04.20$2.8139.0%
- 92026.02.19$9.5036.9%

- 102025.12.17$2.3835.6%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.