IFEval-Ko
KoreanHigher is better
The Korean version of Google's IFEval: instructions with automatically checkable conditions, such as length, keywords, and output format. The score averages four strict and loose accuracy measures. Horangi runs a fixed 100-question subset rather than the full dataset, so its numbers are not directly comparable with full-dataset scores. Scores run from 0 to 100%. Higher is better.
Top score95.0%GPT-5.4
Models tested48
Last updated2026.08.03
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelSettingReleasedPrice/1MScore
- 1xhigh-effort2026.03.05$1295.0%
- 2xhigh-effort2026.04.23$2493.3%
- 32025.08.07$7.8192.9%
- 32026.03.15$3.3092.9%
- 5xhigh-effort2026.03.17$0.9992.5%
- 52026.07.08$592.5%
- 5xhigh-effort2026.03.17$3.5692.5%
- 8high-effort2025.12.17$2.3892.1%

- 9max-effort2026.07.09$1691.7%
- 102026.05.19$7.1391.3%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.