Swallow Leaderboard — Japanese (avg)
Often citedJapaneseHigher is better
The average score of the Japanese leaderboard run by the Swallow project at Institute of Science Tokyo. It averages eight tasks: graduate-level science questions in Japanese (GPQA), Japanese coding (JHumanEval), Japan-specific knowledge (JamC-QA), instruction following (M-IFEval-Ja), multi-domain reasoning (MMLU-ProX), hard math (PolyMath), and English↔Japanese news translation in both directions. Scores run from 0 to 100%. Higher is better.
Top score71.7%GPT-5.4
Models tested8
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelSettingReleasedPrice/1MScore
- 1medium2026.03.05$1271.7%
- 2medium2025.08.07$7.8168.5%
- 32026.04.02$0.3467.5%

- 4medium2025.08.07$1.5663.1%
- 52026.03.11$0.3461.1%
- 62025.04.14$6.5057.1%
- 7medium2025.08.05$0.4056.6%
- 82025.04.05$0.3748.2%
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.