Ko-MTBench
KoreanHigher is better
80 questions, each answered over two turns, across eight areas: writing, roleplay, reasoning, math, coding, extraction, STEM, and humanities. It adapts MT-Bench to Korean language and culture; an AI judge scores answers from 1 to 10, divided by 10. Scores run from 0 to 100%. Higher is better.
Top score83.5%Motif 3
Models tested48
Last updated2026.08.03
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelSettingReleasedPrice/1MScore
- 12026.08.10—83.5%

- 22025.12.17$2.389.0%

- 22026.05.19$7.139.0%

- 2high-effort2025.11.24$209.0%
- 5high-effort2026.02.17$128.9%
- 6high-effort2025.09.29$128.9%
- 7high-effort2026.06.30$88.9%
- 7high-effort2025.08.05$0.408.9%
- 7high-effort2026.02.19$9.508.9%

- 102026.05.21$3.698.9%
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.