Horangi 4 (Korean LLM Leaderboard)
Often citedKoreanHigher is better
The overall score of Horangi 4, W&B Korea's Korean-language LLM leaderboard: 20+ Korean tests are combined into General Language Performance (GLP) and Alignment Performance (ALT), and the score is their average. The same model can appear separately for different reasoning-effort settings. Scores run from 0 to 100%. Higher is better.
Top score84.1%Gemini 3.1 Pro
Models tested49
Last updated2026.08.03
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelSettingReleasedPrice/1MScore
- 1high-effort2026.02.19$9.5084.1%

- 2high-effort2026.06.09$4082.6%
- 3xhigh-effort2026.04.23$2481.8%
- 42026.05.19$7.1381.2%

- 5max-effort2026.07.09$1680.7%
- 62026.05.21$3.6980.2%
- 7xhigh-effort2026.03.05$1279.7%
- 82026.07.08$578.7%
- 9high-effort2026.06.30$878.6%
- 10max-effort2026.07.09$9.5078.3%
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.