MATH-500
In plain words · How many of 500 competition-level math problems it solves
- What does it measure?
- Artificial Analysis now classifies this as 'legacy' and rarely measures new models, so recent ones may be missing. 500 competition-level maths problems covering arithmetic, algebra, geometry and probability. Easier than AIME but broader.
- Who evaluates it, and how?
- A 500-problem representative sample that OpenAI drew in a 2023 study from the MATH dataset (12,500 problems) created by UC Berkeley researchers. Answers are fixed, so grading is automatic; Artificial Analysis runs it independently.
- Reading the score
- 0–100%. Most recent AIs exceed 95%, so it is effectively a race for full marks and cannot separate them on its own.
Rank
Top score99.4%GPT-5
Models tested10
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelReleasedPrice/1MScore
- 12025.08.07$7.8199.4%
- 22025.05.22$1299.1%
- 32025.05.22$6098.2%
- 42025.06.17$1.9598.1%

- 52025.06.17$7.8196.7%

- 62025.04.14$6.5091.3%
- 72025.04.05$0.7088.9%
- 82025.04.05$0.3784.4%
- 92025.01.28$181.7%
- 102026.03.16$0.4956.2%
37 of the 37 AIs released in the last 60 days have no score here yet · Claude Haiku 5.5, Decider V1.1 27B, GPT-6 Luna Decisions and 34 more
Source: Artificial Analysis Data sources & removal requestsWhat to keep in mind
- This score comes from a fixed, predefined evaluation, so it may differ from what you get with your own question.
- Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.