HRM8K

KoreanHigher is better

Math word problems in Korean, combining problems from Korean exams, such as the Korean Mathematical Olympiad and the national college entrance exam (CSAT), with translated problems from well-known English math sets. Final answers are checked by a rule-based grader. Horangi runs a fixed 100-question subset rather than the full dataset, so its numbers are not directly comparable with full-dataset scores. Scores run from 0 to 100%. Higher is better.

Top score97.0%Gemini 3 Flash
Models tested48
Last updated2026.08.03
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    Google2025.12.17 · high-effort
    97.0%
  2. 1
    Anthropic2026.06.09 · high-effort
    97.0%
  3. 1
    Alibaba2026.05.21
    97.0%
  4. 4
    OpenAI2026.03.17 · xhigh-effort
    96.0%
  5. 4
    Google2026.05.19
    96.0%
  6. 4
    OpenAI2026.03.17 · xhigh-effort
    96.0%
  7. 7
    DeepSeek2025.12.01 · high-effort
    95.8%
  8. 8
    Anthropic2026.05.27 · high-effort
    95.0%
  9. 8
    Google2026.02.19 · high-effort
    95.0%
  10. 8
    LG AI Research2026.01.12 · thinking
    95.0%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.