Japanese MT-Bench

JapaneseHigher is better

80 Japanese questions, each answered over two turns, across eight areas: writing, roleplay, reasoning, math, coding, extraction, STEM, and humanities. An AI judge scores answers from 1 to 10, divided by 10. Scores run from 0 to 100%. Higher is better.

Top score98.8%Claude Opus 5
Models tested62
Last updated2026.09.10
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    Anthropic2026.07.24 · adaptive-thinking-max
    98.8%
  2. 2
    Z.ai2026.08.26 · reasoning-max
    98.8%
  3. 3
    Anthropic2026.02.17 · extended-thinking
    98.4%
  4. 4
    Moonshot AI2026.07.17 · reasoning-max
    98.4%
  5. 5
    Google2026.07.21
    98.3%
  6. 6
    Anthropic2026.05.27 · adaptive-thinking-xhigh
    98.1%
  7. 7
    Anthropic2026.04.16 · adaptive-thinking-xhigh
    98.0%
  8. 8
    Z.ai2026.08.18 · reasoning-max
    97.9%
  9. 9
    Anthropic2025.05.22 · extended-thinking
    97.8%
  10. 9
    Google2026.05.19
    97.8%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.