Ko-Moral
KoreanHigher is better
Judge whether a Korean sentence is unethical, answering yes or no. The sentences come from AI Hub's text-ethics verification data, labeled for categories such as censure, hate, discrimination, sexual content, violence, and crime. Horangi runs a fixed 100-question subset rather than the full dataset, so its numbers are not directly comparable with full-dataset scores. Scores run from 0 to 100%. Higher is better.
Top score90.0%Gemini 3 Flash
Models tested48
Last updated2026.08.03
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelSettingReleasedPrice/1MScore
- 12025.12.17$2.3890.0%

- 2xhigh-effort2026.04.23$2488.0%
- 2max-effort2026.07.09$1688.0%
- 4high-effort2026.02.19$9.5087.0%

- 52025.08.07$7.8186.0%
- 6high-effort2026.06.16$3.5784.0%
- 6high-effort2025.09.29$1284.0%
- 8max-effort2026.07.09$0.9583.0%
- 8high-effort2026.05.27$2083.0%
- 82025.08.05$6083.0%
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.