M-IFEval-Ja
JapaneseHigher is better
Tests whether the AI precisely follows Japanese instructions that carry format constraints. It is the Japanese part of M-IFEval, a multilingual instruction-following test. Scores run from 0 to 100%. Higher is better.
Top score90.7%GPT-5
Models tested8
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelSettingReleasedPrice/1MScore
- 1medium2025.08.07$7.8190.7%
- 22026.04.02$0.3489.8%

- 3medium2026.03.05$1288.1%
- 4medium2025.08.07$1.5682.7%
- 52026.03.11$0.3482.3%
- 62025.04.14$6.5081.0%
- 7medium2025.08.05$0.4073.5%
- 82025.04.05$0.3761.1%
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.