MATH-500

In plain words · How many of 500 competition-level math problems it solves

What does it measure?
Artificial Analysis now classifies this as 'legacy' and rarely measures new models, so recent ones may be missing. 500 competition-level maths problems covering arithmetic, algebra, geometry and probability. Easier than AIME but broader.
Who evaluates it, and how?
A 500-problem representative sample that OpenAI drew in a 2023 study from the MATH dataset (12,500 problems) created by UC Berkeley researchers. Answers are fixed, so grading is automatic; Artificial Analysis runs it independently.
Reading the score
0–100%. Most recent AIs exceed 95%, so it is effectively a race for full marks and cannot separate them on its own.

Rank

Top score99.4%GPT-5
Models tested10
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    OpenAI2025.08.07
    99.4%
  2. 2
    Anthropic2025.05.22
    99.1%
  3. 3
    Anthropic2025.05.22
    98.2%
  4. 4
    Google2025.06.17
    98.1%
  5. 5
    Google2025.06.17
    96.7%
  6. 6
    OpenAI2025.04.14
    91.3%
  7. 7
    Meta2025.04.05
    88.9%
  8. 8
    Meta2025.04.05
    84.4%
  9. 9
    Perplexity2025.01.28
    81.7%
  10. 10
    Mistral AI2026.03.16
    56.2%

37 of the 37 AIs released in the last 60 days have no score here yet · Claude Haiku 5.5, Decider V1.1 27B, GPT-6 Luna Decisions and 34 more

Source: Artificial Analysis Data sources & removal requests

What to keep in mind

  • This score comes from a fixed, predefined evaluation, so it may differ from what you get with your own question.
  • Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.