KoBALT-700 (semantics)
KoreanHigher is better
The semantics part of KoBALT-700, a set of Korean linguistics questions written by experts: multiple-choice items on meaning, such as whether a predicate and its arguments fit together semantically. Horangi runs a fixed 100-question subset rather than the full dataset, so its numbers are not directly comparable with full-dataset scores. Scores run from 0 to 100%. Higher is better.
Top score91.0%Gemini 3.1 Pro
Models tested48
Last updated2026.08.03
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelSettingReleasedPrice/1MScore
- 1high-effort2026.02.19$9.5091.0%

- 2high-effort2026.06.09$4087.0%
- 3max-effort2026.07.09$9.5084.0%
- 3max-effort2026.07.09$1684.0%
- 5xhigh-effort2026.04.23$2483.0%
- 62026.05.19$7.1382.8%

- 7high-effort2025.12.17$2.3882.0%

- 72026.02.11$2.4082.0%
- 92026.05.21$3.6981.3%
- 102026.07.08$580.6%
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.