Horangi 4 — ALT (alignment)

KoreanHigher is better

Horangi 4's Alignment Performance (ALT) average: an equal-weight average of five areas in Korean, namely following instructions, ethical judgment, spotting hate speech, avoiding bias, and avoiding hallucination. It covers instruction-following and hallucination, not just safety. Scores run from 0 to 100%. Higher is better.

Top score84.8%Gemini 3.1 Pro
Models tested48
Last updated2026.08.03
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    Google2026.02.19 · high-effort
    84.8%
  2. 2
    Alibaba2026.05.21
    81.6%
  3. 3
    Anthropic2026.06.09 · high-effort
    81.3%
  4. 4
    Google2025.12.17
    81.2%
  5. 5
    Z.ai2026.06.16 · high-effort
    81.1%
  6. 6
    OpenAI2026.04.23 · xhigh-effort
    80.8%
  7. 7
    Z.ai2026.03.15
    80.4%
  8. 8
    Google2026.05.19
    80.1%
  9. 9
    Anthropic2025.08.05
    80.0%
  10. 10
    Anthropic2025.05.22 · high-effort
    79.6%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.