Japanese MT-Bench (Swallow)

JapaneseHigher is better

Japanese MT-Bench as run by the Swallow leaderboard itself. Models answer questions over two turns across eight areas (writing, roleplay, reasoning, math, coding, extraction, STEM, humanities) and an AI judge grades them. The questions come from the Nejumi edition but grading is done separately, so values differ from Nejumi's Japanese MT-Bench score. Scores run from 0 to 100%. Higher is better.

Top score84.4%GPT-5.4
Models tested8
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    OpenAI2026.03.05 · medium
    84.4%
  2. 2
    OpenAI2025.08.07 · medium
    84.2%
  3. 3
    OpenAI2025.08.07 · medium
    83.0%
  4. 4
    Google2026.04.02
    81.5%
  5. 5
    OpenAI2025.04.14
    77.1%
  6. 6
    OpenAI2025.08.05 · medium
    75.7%
  7. 7
    NVIDIA2026.03.11
    73.6%
  8. 8
    Meta2025.04.05
    62.9%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.