Ko-Moral

KoreanHigher is better

Judge whether a Korean sentence is unethical, answering yes or no. The sentences come from AI Hub's text-ethics verification data, labeled for categories such as censure, hate, discrimination, sexual content, violence, and crime. Horangi runs a fixed 100-question subset rather than the full dataset, so its numbers are not directly comparable with full-dataset scores. Scores run from 0 to 100%. Higher is better.

Top score90.0%Gemini 3 Flash
Models tested48
Last updated2026.08.03
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    Google2025.12.17
    90.0%
  2. 2
    OpenAI2026.04.23 · xhigh-effort
    88.0%
  3. 2
    OpenAI2026.07.09 · max-effort
    88.0%
  4. 4
    Google2026.02.19 · high-effort
    87.0%
  5. 5
    OpenAI2025.08.07
    86.0%
  6. 6
    Z.ai2026.06.16 · high-effort
    84.0%
  7. 6
    Anthropic2025.09.29 · high-effort
    84.0%
  8. 8
    OpenAI2026.07.09 · max-effort
    83.0%
  9. 8
    Anthropic2026.05.27 · high-effort
    83.0%
  10. 8
    Anthropic2025.08.05
    83.0%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.