Japanese MT-Bench
JapaneseHigher is better
80 Japanese questions, each answered over two turns, across eight areas: writing, roleplay, reasoning, math, coding, extraction, STEM, and humanities. An AI judge scores answers from 1 to 10, divided by 10. Scores run from 0 to 100%. Higher is better.
Top score98.8%Claude Opus 5
Models tested62
Last updated2026.09.10
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelSettingReleasedPrice/1MScore
- 1adaptive-thinking-max2026.07.24$2098.8%
- 2reasoning-max2026.08.26$0.4198.8%
- 3extended-thinking2026.02.17$1298.4%
- 4reasoning-max2026.07.17$1198.4%
- 52026.07.21$398.3%

- 6adaptive-thinking-xhigh2026.05.27$2098.1%
- 7adaptive-thinking-xhigh2026.04.16$2098.0%
- 8reasoning-max2026.08.18$3.6097.9%
- 9extended-thinking2025.05.22$1297.8%
- 92026.05.19$7.1397.8%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.