Arena Japanese

In plain words · Which AI people preferred when asked in Japanese

What does it measure?
The ranking computed only on prompts written in Japanese. It can expose AIs that score well in English but fall short in Japanese expression and context.
Who evaluates it, and how?
On Arena Intelligence, a user enters a question, two anonymous AIs answer, and the user votes for the better one. Over 82 million votes are aggregated with a statistical model (Bradley-Terry) into an Elo-style score, and this table takes the 'style-controlled' version from the official dataset, which statistically removes the bias toward long, nicely formatted answers. The prompt language is detected automatically and only Japanese prompts are counted.
Reading the score
This is a head-to-head record, not an exam score (accuracy %). It is computed from the probability of being preferred over other AIs, so there is no maximum, and when a new AI enters, existing scores shift slightly. Typical values are 1,000–1,500; a 100-point gap means the leader is preferred about 64% of the time, a 200-point gap about 76%. Correctness is not counted — a wrong but friendly answer can win — so if accuracy matters, also look at accuracy tests (GPQA, AIME, etc.). Gaps of a few dozen points are effectively the same tier even if the rank differs. The sample is smaller than the English overall ranking, so ranks move more, and Japanese rankings correlate weakly with other languages — treat them separately.

Last updated: 2026-10-02

Rank

Top score1542Claude Fable 5.1
Models tested75
Last updated2026.10.02
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    Anthropic2026.09.01
    1542
  2. 2
    Anthropic2026.06.09
    1524
  3. 3
    Google2026.08.13
    1504
  4. 4
    Google2026.09.03
    1503
  5. 5
    Moonshot AI2026.07.17
    1502
  6. 6
    Google2026.02.19
    1499
  7. 7
    Meta2026.09.02
    1498
  8. 8
    Anthropic2026.07.24
    1497
  9. 9
    OpenAI2026.07.09
    1495
  10. 10
    Google2025.12.17
    1487

28 of the 37 AIs released in the last 60 days have no score here yet · Claude Haiku 5.5, Decider V1.1 27B, Mistral Large 4 and 25 more

Source: Arena Intelligence Data sources & removal requests

What to keep in mind

  • It’s a preference vote, not a correctness check. A wrong but friendly answer can win.
  • Which questions got asked depends on who voted — they may differ from your use.
  • Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.