Ko-HLE
KoreanHigher is better
A Korean translation of Humanity's Last Exam, expert-level questions across many fields, answered without tools. It is not comparable with vendor-reported HLE scores, which use the full 2,500 questions and often allow tools. Horangi runs a fixed 100-question subset rather than the full dataset, so its numbers are not directly comparable with full-dataset scores. Scores run from 0 to 100%. Higher is better.
Top score53.1%Claude Fable 5
Models tested48
Last updated2026.08.03
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelSettingReleasedPrice/1MScore
- 1high-effort2026.06.09$4053.1%
- 2max-effort2026.07.09$1652.0%
- 32026.07.08$550.0%
- 4max-effort2026.07.09$9.5047.0%
- 4xhigh-effort2026.03.05$1247.0%
- 6xhigh-effort2026.04.23$2446.5%
- 72026.05.19$7.1345.0%

- 7high-effort2026.02.19$9.5045.0%

- 92026.05.21$3.6942.4%
- 102025.12.17$2.3841.0%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.