AIME 2026

Widely citedHigher is better

The 30 problems of the February 2026 AIME (15 each in AIME I and II), the qualifying exam for the US math olympiad; every answer is an integer from 0 to 999, and top models are near 100%. MathArena runs each model four times per problem, right after the contest, and averages the accuracy. Scores run from 0 to 100%. Higher is better.

Top score100.0%GPT-5.5
Models tested16
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    OpenAI2026.04.23 · xhigh
    100.0%
  2. 1
    Anthropic2026.05.27 · max
    100.0%
  3. 3
    OpenAI2026.03.05 · xhigh
    99.2%
  4. 4
    Google2026.02.19
    98.3%
  5. 5
    Google2025.12.17
    96.7%
  6. 5
    Anthropic2026.02.04 · high
    96.7%
  7. 5
    Google2026.07.21
    96.7%
  8. 5
    Z.ai2026.02.11
    96.7%
  9. 5
    DeepSeek2026.04.24 · max
    96.7%
  10. 10
    Anthropic2026.04.16 · xhigh
    95.8%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.