Swallow Leaderboard — Japanese (avg)

Often citedJapaneseHigher is better

The average score of the Japanese leaderboard run by the Swallow project at Institute of Science Tokyo. It averages eight tasks: graduate-level science questions in Japanese (GPQA), Japanese coding (JHumanEval), Japan-specific knowledge (JamC-QA), instruction following (M-IFEval-Ja), multi-domain reasoning (MMLU-ProX), hard math (PolyMath), and English↔Japanese news translation in both directions. Scores run from 0 to 100%. Higher is better.

Top score71.7%GPT-5.4
Models tested8
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    OpenAI2026.03.05 · medium
    71.7%
  2. 2
    OpenAI2025.08.07 · medium
    68.5%
  3. 3
    Google2026.04.02
    67.5%
  4. 4
    OpenAI2025.08.07 · medium
    63.1%
  5. 5
    NVIDIA2026.03.11
    61.1%
  6. 6
    OpenAI2025.04.14
    57.1%
  7. 7
    OpenAI2025.08.05 · medium
    56.6%
  8. 8
    Meta2025.04.05
    48.2%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.