AIME 2026
Widely citedHigher is better
The 30 problems of the February 2026 AIME (15 each in AIME I and II), the qualifying exam for the US math olympiad; every answer is an integer from 0 to 999, and top models are near 100%. MathArena runs each model four times per problem, right after the contest, and averages the accuracy. Scores run from 0 to 100%. Higher is better.
Top score100.0%GPT-5.5
Models tested16
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelSettingReleasedPrice/1MScore
- 1xhigh2026.04.23$24100.0%
- 1max2026.05.27$20100.0%
- 3xhigh2026.03.05$1299.2%
- 42026.02.19$9.5098.3%

- 52025.12.17$2.3896.7%

- 5high2026.02.04$2096.7%
- 52026.07.21$396.7%

- 52026.02.11$2.4096.7%
- 5max2026.04.24$2.7396.7%
- 10xhigh2026.04.16$2095.8%
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.