MMLU-Pro

In plain words · How many college-level questions across 14 subjects it gets right

What does it measure?
Artificial Analysis now classifies this as 'legacy' and rarely measures new models, so recent ones may be missing. A harder version of MMLU, the exam built from university-level subjects. It tests knowledge and reasoning across a broad range of fields — law, medicine, engineering, history and more.
Who evaluates it, and how?
Researchers at the University of Waterloo and elsewhere expanded MMLU from four to ten answer options, removed questions that were too easy or flawed, and rebuilt it into about 12,000 questions across 14 fields. Scored as accuracy; Artificial Analysis runs it independently.
Reading the score
0–100%. With ten options, guessing gives 10%. It is designed to score 16–33 points lower than the original MMLU, so do not compare it directly with MMLU numbers elsewhere.

Rank

Top score89.5%Claude Opus 4.5
Models tested28
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 🥇AnthropicClaude Opus 4.5Reasoning effort: Thinking89.5%
  2. 🥈GoogleGemini 3 FlashReasoning effort: Thinking89.0%
  3. 🥉AnthropicClaude Opus 4.1Reasoning effort: Thinking88.0%
  4. #4AnthropicClaude Sonnet 4.5Reasoning effort: Thinking87.5%
  5. #5AnthropicClaude Opus 4Reasoning effort: Thinking87.3%
  6. #6OpenAIGPT-5Reasoning effort: High87.1%
  7. #7DeepSeekDeepSeek V3.2Reasoning effort: Thinking86.2%
  8. #7GoogleGemini 2.5 Pro86.2%
  9. #7UpstageSolar Pro 486.2%
  10. #10xAIGrok 4.1 Fast (Reasoning)Reasoning effort: Thinking85.4%
  11. #11AnthropicClaude Sonnet 4Reasoning effort: Thinking84.2%
  12. #12LG AI ResearchK-EXAONEReasoning effort: Thinking83.8%
  13. #13OpenAIGPT-5 MiniReasoning effort: High83.7%
  14. #14LG AI ResearchK-EXAONE 2.083.5%
  15. #15GoogleGemini 2.5 FlashReasoning effort: Thinking83.2%
  16. #16BaiduERNIE 5.0 Thinking83.0%
  17. #17AmazonNova 2 LiteReasoning effort: High81.8%
  18. #18MetaLlama 4 Maverick80.9%
  19. #19GoogleGemini 2.5 Flash LiteReasoning effort: Thinking80.8%
  20. #19OpenAIGPT OSS 120BReasoning effort: High80.8%
  21. #21OpenAIGPT-4.180.6%
  22. #22AnthropicClaude Haiku 4.5Reasoning effort: None80.0%
  23. #23NVIDIANemotron 3 Nano 30B A3BReasoning effort: Thinking79.4%
  24. #24OpenAIGPT-5 NanoReasoning effort: High78.0%
  25. #25MetaLlama 4 Scout75.2%
  26. #26xAIGrok 4.1 FastReasoning effort: None74.3%
  27. #27PerplexitySonar68.9%
  28. #28Mistral AIMistral Small 441.9%
Half of models ≤ 83.3%

37 of the 37 AIs released in the last 60 days have no score here yet · Claude Haiku 5.5, Decider V1.1 27B, GPT-6 Luna Decisions and 34 more

Source: Artificial Analysis Data sources & removal requests

Related news

What to keep in mind

  • This score comes from a fixed, predefined evaluation, so it may differ from what you get with your own question.
  • Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.