HAE-RAE Bench v1 (Reading)
KoreanHigher is better
Only the reading-comprehension part of HAE-RAE Bench 1.0: Korean multiple-choice questions taken from the Korean Language Ability Test (KLAT). The other areas, such as standard terms, loanwords, rare words, general knowledge, and history, are not included. Horangi runs a fixed 100-question subset rather than the full dataset, so its numbers are not directly comparable with full-dataset scores. Scores run from 0 to 100%. Higher is better.
Top score93.0%GPT-5.6 Luna
Models tested48
Last updated2026.08.03
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelSettingReleasedPrice/1MScore
- 1max-effort2026.07.09$0.9593.0%
- 1high-effort2026.06.09$4093.0%
- 12026.05.19$7.1393.0%

- 1max-effort2026.07.09$1693.0%
- 5high-effort2025.12.17$2.3892.0%

- 5xhigh-effort2026.04.23$2492.0%
- 5high-effort2026.02.17$1292.0%
- 5high-effort2026.05.27$2092.0%
- 52025.09.15$0.4292.0%
- 10high-effort2026.06.30$891.0%
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.