KMMLU-Pro
Often citedKoreanHigher is better
Multiple-choice questions from 14 Korean national professional licensing exams, such as medicine, law, tax, and accounting. The paper also reports pass/fail per license, but the value here is plain accuracy. Horangi runs a fixed 100-question subset rather than the full dataset, so its numbers are not directly comparable with full-dataset scores. Scores run from 0 to 100%. Higher is better.
Top score95.0%Claude Opus 4.8
Models tested48
Last updated2026.08.03
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelSettingReleasedPrice/1MScore
- 1high-effort2026.05.27$2095.0%
- 1max-effort2026.07.09$1695.0%
- 3xhigh-effort2026.04.23$2494.0%
- 4high-effort2025.12.17$2.3893.0%

- 5max-effort2026.07.09$9.5091.0%
- 5xhigh-effort2026.03.05$1291.0%
- 7high-effort2026.06.30$890.0%
- 7max-effort2026.07.09$0.9590.0%
- 92026.05.19$7.1389.9%

- 9high-effort2026.02.19$9.5089.9%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.