Korean Hate Speech

KoreanHigher is better

Judge whether comments from a Korean entertainment-news site are hate speech. It measures how accurately the model detects hate speech, not whether it produces any. Horangi runs a fixed 100-question subset rather than the full dataset, so its numbers are not directly comparable with full-dataset scores. Scores run from 0 to 100%. Higher is better.

Top score78.0%Claude Fable 5
Models tested48
Last updated2026.08.03
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    Anthropic2026.06.09 · high-effort
    78.0%
  2. 2
    Anthropic2025.10.15 · high-effort
    76.0%
  3. 3
    Anthropic2025.05.22 · high-effort
    71.0%
  4. 4
    DeepSeek2025.12.01 · high-effort
    70.0%
  5. 4
    xAI2025.11.19
    70.0%
  6. 6
    Mistral AI2026.04.28
    69.0%
  7. 6
    Z.ai2026.03.15
    69.0%
  8. 8
    Anthropic2026.02.17 · high-effort
    67.0%
  9. 8
    Alibaba2026.02.16
    67.0%
  10. 8
    Mistral AI2026.03.16
    67.0%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.