Horangi 4 — GLP (general language)

KoreanHigher is better

Horangi 4's General Language Performance (GLP) average: a weighted average of 11 areas, including Korean grammar and meaning, expression, knowledge, and math and logical reasoning, plus coding and function calling. Those last two are English and code tasks, so it is not a pure measure of Korean ability. Scores run from 0 to 100%. Higher is better.

Top score83.9%Claude Fable 5
Models tested49
Last updated2026.08.03
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    Anthropic2026.06.09 · high-effort
    83.9%
  2. 2
    Google2026.02.19 · high-effort
    83.4%
  3. 3
    OpenAI2026.04.23 · xhigh-effort
    82.8%
  4. 4
    OpenAI2026.07.09 · max-effort
    82.7%
  5. 5
    Google2026.05.19
    82.3%
  6. 6
    OpenAI2026.03.05 · xhigh-effort
    81.9%
  7. 7
    OpenAI2026.07.09 · max-effort
    81.6%
  8. 8
    xAI2026.07.08
    80.2%
  9. 9
    Anthropic2026.06.30 · high-effort
    79.3%
  10. 10
    Alibaba2026.05.21
    78.8%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.