IFEval-Ko

KoreanHigher is better

The Korean version of Google's IFEval: instructions with automatically checkable conditions, such as length, keywords, and output format. The score averages four strict and loose accuracy measures. Horangi runs a fixed 100-question subset rather than the full dataset, so its numbers are not directly comparable with full-dataset scores. Scores run from 0 to 100%. Higher is better.

Top score95.0%GPT-5.4
Models tested48
Last updated2026.08.03
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    OpenAI2026.03.05 · xhigh-effort
    95.0%
  2. 2
    OpenAI2026.04.23 · xhigh-effort
    93.3%
  3. 3
    OpenAI2025.08.07
    92.9%
  4. 3
    Z.ai2026.03.15
    92.9%
  5. 5
    OpenAI2026.03.17 · xhigh-effort
    92.5%
  6. 5
    xAI2026.07.08
    92.5%
  7. 5
    OpenAI2026.03.17 · xhigh-effort
    92.5%
  8. 8
    Google2025.12.17 · high-effort
    92.1%
  9. 9
    OpenAI2026.07.09 · max-effort
    91.7%
  10. 10
    Google2026.05.19
    91.3%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.