Horangi 4 (Korean LLM Leaderboard)

Often citedKoreanHigher is better

The overall score of Horangi 4, W&B Korea's Korean-language LLM leaderboard: 20+ Korean tests are combined into General Language Performance (GLP) and Alignment Performance (ALT), and the score is their average. The same model can appear separately for different reasoning-effort settings. Scores run from 0 to 100%. Higher is better.

Top score84.1%Gemini 3.1 Pro
Models tested49
Last updated2026.08.03
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    Google2026.02.19 · high-effort
    84.1%
  2. 2
    Anthropic2026.06.09 · high-effort
    82.6%
  3. 3
    OpenAI2026.04.23 · xhigh-effort
    81.8%
  4. 4
    Google2026.05.19
    81.2%
  5. 5
    OpenAI2026.07.09 · max-effort
    80.7%
  6. 6
    Alibaba2026.05.21
    80.2%
  7. 7
    OpenAI2026.03.05 · xhigh-effort
    79.7%
  8. 8
    xAI2026.07.08
    78.7%
  9. 9
    Anthropic2026.06.30 · high-effort
    78.6%
  10. 10
    OpenAI2026.07.09 · max-effort
    78.3%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.