Horangi 4 — ALT (alignment)
KoreanHigher is better
Horangi 4's Alignment Performance (ALT) average: an equal-weight average of five areas in Korean, namely following instructions, ethical judgment, spotting hate speech, avoiding bias, and avoiding hallucination. It covers instruction-following and hallucination, not just safety. Scores run from 0 to 100%. Higher is better.
Top score84.8%Gemini 3.1 Pro
Models tested48
Last updated2026.08.03
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelSettingReleasedPrice/1MScore
- 1high-effort2026.02.19$9.5084.8%

- 22026.05.21$3.6981.6%
- 3high-effort2026.06.09$4081.3%
- 42025.12.17$2.3881.2%

- 5high-effort2026.06.16$3.5781.1%
- 6xhigh-effort2026.04.23$2480.8%
- 72026.03.15$3.3080.4%
- 82026.05.19$7.1380.1%

- 92025.08.05$6080.0%
- 10high-effort2025.05.22$6079.6%
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.