Nejumi 4 — GLP (general language)
JapaneseHigher is better
Nejumi 4's General Language Performance (GLP) average across five broad areas: Japanese expression, translation, and question answering; reasoning; knowledge; grammar and meaning; and coding and function calling. Scores run from 0 to 100%. Higher is better.
Top score86.7%Claude Opus 5
Models tested62
Last updated2026.09.10
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelSettingReleasedPrice/1MScore
- 1adaptive-thinking-max2026.07.24$2086.7%
- 2adaptive-thinking-max with fallback to Opus 4.82026.06.09$4085.0%
- 3max-effort2026.07.09$1684.0%
- 4max-effort2026.07.09$9.5083.6%
- 5adaptive-thinking-xhigh2026.05.27$2083.3%
- 6high-effort2026.04.23$2483.1%
- 7adaptive-thinking-xhigh2026.04.16$2083.0%
- 82026.02.19$9.5082.8%

- 92026.07.21$382.5%

- 10reasoning-max2026.07.17$1182.0%
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.