Ko-HLE

KoreanHigher is better

A Korean translation of Humanity's Last Exam, expert-level questions across many fields, answered without tools. It is not comparable with vendor-reported HLE scores, which use the full 2,500 questions and often allow tools. Horangi runs a fixed 100-question subset rather than the full dataset, so its numbers are not directly comparable with full-dataset scores. Scores run from 0 to 100%. Higher is better.

Top score53.1%Claude Fable 5
Models tested48
Last updated2026.08.03
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    Anthropic2026.06.09 · high-effort
    53.1%
  2. 2
    OpenAI2026.07.09 · max-effort
    52.0%
  3. 3
    xAI2026.07.08
    50.0%
  4. 4
    OpenAI2026.07.09 · max-effort
    47.0%
  5. 4
    OpenAI2026.03.05 · xhigh-effort
    47.0%
  6. 6
    OpenAI2026.04.23 · xhigh-effort
    46.5%
  7. 7
    Google2026.05.19
    45.0%
  8. 7
    Google2026.02.19 · high-effort
    45.0%
  9. 9
    Alibaba2026.05.21
    42.4%
  10. 10
    Google2025.12.17
    41.0%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.