AIME 2025
In plain words · How many problems from a top US high-school math contest it solves
- What does it measure?
- Artificial Analysis now classifies this as 'legacy' and rarely measures new models, so recent ones may be missing. The AI solves the actual 2025 problems from AIME, the qualifying round of the US high-school mathematics olympiad. It tests mathematical thinking — working an unfamiliar problem through to the end — rather than knowing formulas.
- Who evaluates it, and how?
- All 30 problems from the 2025 AIME set by the Mathematical Association of America. Answers are integers from 0 to 999, so grading is automatic; Artificial Analysis runs it independently.
- Reading the score
- 0–100%. Perfect scores are very rare for humans, but recent reasoning models often exceed 90%, so the test has quickly lost its ability to separate them. With only 30 problems, one problem is about 3.3 points.
Rank
Top score97.0%Gemini 3 Flash
Models tested24
Model release date (newest on the right)
Best score so farHigher on the chart is better
Can it solve problems that stump experts across fields?Humanity's Last Exam
Claude Opus 5.5🥇
Claude Opus 5.5🥇Can it solve research-level physics problems?CritPt
GPT-5.6 Sol🥇
GPT-5.6 Sol🥇Which AI gives the math answers people prefer most?Arena Math
Gemini 4 Argon🥇
Gemini 4 Argon🥇
- 🥇GoogleGemini 3 FlashReasoning effort: Thinking97.0%
- 🥈OpenAIGPT-5Reasoning effort: High94.3%
- 🥈AmazonNova 2 LiteReasoning effort: High94.3%
- #4OpenAIGPT OSS 120BReasoning effort: High93.4%
- #5DeepSeekDeepSeek V3.2Reasoning effort: Thinking92.0%
- #6AnthropicClaude Opus 4.5Reasoning effort: Thinking91.3%
- #7NVIDIANemotron 3 Nano 30B A3BReasoning effort: Thinking91.0%
- #8OpenAIGPT-5 MiniReasoning effort: High90.7%
- #9LG AI ResearchK-EXAONEReasoning effort: Thinking90.3%
- #10xAIGrok 4.1 Fast (Reasoning)Reasoning effort: Thinking89.3%
- #11AnthropicClaude Sonnet 4.5Reasoning effort: Thinking88.0%
- #12GoogleGemini 2.5 Pro87.7%
- #13BaiduERNIE 5.0 Thinking85.0%
- #14AnthropicClaude Haiku 4.5Reasoning effort: Thinking83.7%
- #14OpenAIGPT-5 NanoReasoning effort: High83.7%
- #16AnthropicClaude Opus 4.1Reasoning effort: Thinking80.3%
- #17AnthropicClaude Sonnet 4Reasoning effort: Thinking74.3%
- #18AnthropicClaude Opus 4Reasoning effort: Thinking73.3%
- #18GoogleGemini 2.5 FlashReasoning effort: Thinking73.3%
- #20GoogleGemini 2.5 Flash LiteReasoning effort: Thinking68.7%
- #21OpenAIGPT-4.134.7%
- #22xAIGrok 4.1 FastReasoning effort: None34.3%
- #23MetaLlama 4 Maverick19.3%
- #24MetaLlama 4 Scout14.0%
Half of models ≤ 86.3%
37 of the 37 AIs released in the last 60 days have no score here yet · Claude Haiku 5.5, Decider V1.1 27B, GPT-6 Luna Decisions and 34 more
Source: Artificial Analysis Data sources & removal requestsWhat to keep in mind
- This score comes from a fixed, predefined evaluation, so it may differ from what you get with your own question.
- Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.