M-IFEval-Ja

JapaneseHigher is better

Tests whether the AI precisely follows Japanese instructions that carry format constraints. It is the Japanese part of M-IFEval, a multilingual instruction-following test. Scores run from 0 to 100%. Higher is better.

Top score90.7%GPT-5
Models tested8
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    OpenAI2025.08.07 · medium
    90.7%
  2. 2
    Google2026.04.02
    89.8%
  3. 3
    OpenAI2026.03.05 · medium
    88.1%
  4. 4
    OpenAI2025.08.07 · medium
    82.7%
  5. 5
    NVIDIA2026.03.11
    82.3%
  6. 6
    OpenAI2025.04.14
    81.0%
  7. 7
    OpenAI2025.08.05 · medium
    73.5%
  8. 8
    Meta2025.04.05
    61.1%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.