PolyMath (Japanese, HT)

JapaneseHigher is better

The Japanese edition of PolyMath, a multilingual test of hard math. The model must read and solve difficult problems written in Japanese. This is the variant (HT) used by the Swallow leaderboard. Scores run from 0 to 100%. Higher is better.

Top score77.6%GPT-5.4
Models tested8
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    OpenAI2026.03.05 · medium
    77.6%
  2. 2
    Google2026.04.02
    68.4%
  3. 3
    OpenAI2025.08.07 · medium
    62.4%
  4. 4
    NVIDIA2026.03.11
    54.4%
  5. 5
    OpenAI2025.08.07 · medium
    51.2%
  6. 6
    OpenAI2025.08.05 · medium
    46.4%
  7. 7
    OpenAI2025.04.14
    14.8%
  8. 7
    Meta2025.04.05
    14.8%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.