Ko-ARC-AGI
KoreanHigher is better
ARC-AGI grid puzzles (infer a transformation rule from examples and draw the output for a new input) given with Korean instructions. The output grid must match exactly in size and values. Scores are not comparable with official ARC Prize results. Horangi runs a fixed 100-question subset rather than the full dataset, so its numbers are not directly comparable with full-dataset scores. Scores run from 0 to 100%. Higher is better.
Top score97.0%GPT-5.6 Terra
Models tested48
Last updated2026.08.03
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelSettingReleasedPrice/1MScore
- 1max-effort2026.07.09$9.5097.0%
- 1max-effort2026.07.09$1697.0%
- 3high-effort2026.06.09$4096.0%
- 4xhigh-effort2026.03.05$1295.0%
- 5xhigh-effort2026.04.23$2493.0%
- 6high-effort2026.02.19$9.5092.0%

- 72026.08.10—90.9%

- 82026.05.19$7.1389.9%

- 9high-effort2026.06.30$885.0%
- 102026.07.08$584.0%
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.