Nejumi 4 — GLP (general language)

JapaneseHigher is better

Nejumi 4's General Language Performance (GLP) average across five broad areas: Japanese expression, translation, and question answering; reasoning; knowledge; grammar and meaning; and coding and function calling. Scores run from 0 to 100%. Higher is better.

Top score86.7%Claude Opus 5
Models tested62
Last updated2026.09.10
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    Anthropic2026.07.24 · adaptive-thinking-max
    86.7%
  2. 2
    Anthropic2026.06.09 · adaptive-thinking-max with fallback to Opus 4.8
    85.0%
  3. 3
    OpenAI2026.07.09 · max-effort
    84.0%
  4. 4
    OpenAI2026.07.09 · max-effort
    83.6%
  5. 5
    Anthropic2026.05.27 · adaptive-thinking-xhigh
    83.3%
  6. 6
    OpenAI2026.04.23 · high-effort
    83.1%
  7. 7
    Anthropic2026.04.16 · adaptive-thinking-xhigh
    83.0%
  8. 8
    Google2026.02.19
    82.8%
  9. 9
    Google2026.07.21
    82.5%
  10. 10
    Moonshot AI2026.07.17 · reasoning-max
    82.0%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.