Global-MMLU-Lite

Often citedHigher is better

MMLU four-option knowledge questions translated or post-edited by people into many languages, with 200 questions needing cultural knowledge and 200 not needing it per language. The value here is Kaggle's average over 16 languages. Scores run from 0 to 100%. Higher is better.

Top score95.7%Claude Opus 5
Models tested23
Last updated2026.09.30
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    Anthropic2026.07.24
    95.7%
  2. 2
    Google2026.05.19
    95.4%
  3. 3
    Anthropic2026.09.22
    95.4%
  4. 4
    Google2026.09.03
    95.3%
  5. 5
    Google2026.08.13
    94.9%
  6. 6
    Anthropic2026.05.27
    94.9%
  7. 7
    Anthropic2025.08.05
    94.3%
  8. 8
    Google2026.07.21
    94.3%
  9. 9
    OpenAI2026.07.09
    94.1%
  10. 10
    OpenAI2026.09.22
    94.1%