Ko-MTBench

KoreanHigher is better

80 questions, each answered over two turns, across eight areas: writing, roleplay, reasoning, math, coding, extraction, STEM, and humanities. It adapts MT-Bench to Korean language and culture; an AI judge scores answers from 1 to 10, divided by 10. Scores run from 0 to 100%. Higher is better.

Top score83.5%Motif 3
Models tested48
Last updated2026.08.03
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    Motif Technologies2026.08.10
    83.5%
  2. 2
    Google2025.12.17
    9.0%
  3. 2
    Google2026.05.19
    9.0%
  4. 2
    Anthropic2025.11.24 · high-effort
    9.0%
  5. 5
    Anthropic2026.02.17 · high-effort
    8.9%
  6. 6
    Anthropic2025.09.29 · high-effort
    8.9%
  7. 7
    Anthropic2026.06.30 · high-effort
    8.9%
  8. 7
    OpenAI2025.08.05 · high-effort
    8.9%
  9. 7
    Google2026.02.19 · high-effort
    8.9%
  10. 10
    Alibaba2026.05.21
    8.9%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.