HAE-RAE Bench v1 (Reading)

KoreanHigher is better

Only the reading-comprehension part of HAE-RAE Bench 1.0: Korean multiple-choice questions taken from the Korean Language Ability Test (KLAT). The other areas, such as standard terms, loanwords, rare words, general knowledge, and history, are not included. Horangi runs a fixed 100-question subset rather than the full dataset, so its numbers are not directly comparable with full-dataset scores. Scores run from 0 to 100%. Higher is better.

Top score93.0%GPT-5.6 Luna
Models tested48
Last updated2026.08.03
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    OpenAI2026.07.09 · max-effort
    93.0%
  2. 1
    Anthropic2026.06.09 · high-effort
    93.0%
  3. 1
    Google2026.05.19
    93.0%
  4. 1
    OpenAI2026.07.09 · max-effort
    93.0%
  5. 5
    Google2025.12.17 · high-effort
    92.0%
  6. 5
    OpenAI2026.04.23 · xhigh-effort
    92.0%
  7. 5
    Anthropic2026.02.17 · high-effort
    92.0%
  8. 5
    Anthropic2026.05.27 · high-effort
    92.0%
  9. 592.0%
  10. 10
    Anthropic2026.06.30 · high-effort
    91.0%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.