KMMLU
KoreanHigher is better
Expert-level Korean multiple-choice questions across 45 subjects, from the humanities to STEM, collected from real Korean exams rather than translated from the English MMLU. Horangi runs a fixed 100-question subset rather than the full dataset, so its numbers are not directly comparable with full-dataset scores. Scores run from 0 to 100%. Higher is better.
Top score94.0%Gemini 3.1 Pro
Models tested48
Last updated2026.08.03
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelSettingReleasedPrice/1MScore
- 1high-effort2026.02.19$9.5094.0%

- 1max-effort2026.07.09$1694.0%
- 32026.05.19$7.1393.9%

- 4max-effort2026.07.09$9.5093.0%
- 5xhigh-effort2026.04.23$2492.0%
- 62026.02.11$2.4091.7%
- 7high-effort2026.02.17$1291.0%
- 7high-effort2025.11.24$2091.0%
- 9high-effort2025.12.17$2.3890.0%

- 9max-effort2026.07.09$0.9590.0%
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.