CMIMC 2025

Higher is better

40 final-answer problems from CMIMC 2025, the Carnegie Mellon Informatics and Mathematics Competition for high-school students. MathArena runs each model four times per problem, right after the contest, and averages the accuracy. Scores run from 0 to 100%. Higher is better.

Top score92.5%GLM-5
Models tested8
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    Z.ai2026.02.11
    92.5%
  2. 2
    Google2025.12.17
    90.6%
  3. 3
    OpenAI2025.08.07 · high
    90.0%
  4. 4
    OpenAI2025.08.05 · high
    85.6%
  5. 5
    OpenAI2025.08.07 · high
    84.4%
  6. 6
    OpenAI2025.08.07 · high
    73.8%
  7. 7
    Google2025.06.17
    58.1%
  8. 8
    Google2025.06.17 · thinking
    51.9%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.