Arena Legal & Government

In plain words · Which AI people preferred on expert-level legal and government questions

What does it measure?
Which AI's answers were preferred on expert-level questions in Legal & Government. Examples: reviewing contract clauses, comparing precedents, interpreting regulations. These are questions the people who actually do that work would ask, not general ones.
Who evaluates it, and how?
On Arena Intelligence, a user enters a question, two anonymous AIs answer, and the user votes for the better one. Over 82 million votes are aggregated with a statistical model (Bradley-Terry) into an Elo-style score, and this table takes the 'style-controlled' version from the official dataset, which statistically removes the bias toward long, nicely formatted answers. An AI classifier judges each prompt's reasoning depth and expertise and tags only about 5.5% as 'expert' prompts, which are then divided into occupational categories based on the US Bureau of Labor Statistics classification (SOC), launched November 2025. This table counts only prompts classified as 'Legal & Government'.
Reading the score
This is a head-to-head record, not an exam score (accuracy %). It is computed from the probability of being preferred over other AIs, so there is no maximum, and when a new AI enters, existing scores shift slightly. Typical values are 1,000–1,500; a 100-point gap means the leader is preferred about 64% of the time, a 200-point gap about 76%. Correctness is not counted — a wrong but friendly answer can win — so if accuracy matters, also look at accuracy tests (GPQA, AIME, etc.). Gaps of a few dozen points are effectively the same tier even if the rank differs. This category has few prompts, so top ranks change often; look at score bands rather than ranks.

Last updated: 2026-10-02

Rank

Top score1537Muse Spark 1.2
Models tested97
Last updated2026.10.02
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    Meta2026.08.05
    1537
  2. 2
    Google2026.09.30
    1533
  3. 3
    Moonshot AI2026.07.17
    1512
  4. 4
    OpenAI2026.09.29
    1511
  5. 5
    Anthropic2026.06.09
    1511
  6. 6
    Anthropic2026.02.04
    1506
  7. 7
    Meta2026.09.02
    1504
  8. 8
    Anthropic2026.07.24
    1503
  9. 9
    Meta2026.04.08
    1501
  10. 10
    Google2026.02.19
    1500

18 of the 37 AIs released in the last 60 days have no score here yet · Claude Haiku 5.5, Decider V1.1 27B, GPT-6 Luna Decisions and 15 more

Source: Arena Intelligence Data sources & removal requests

What to keep in mind

  • It’s a preference vote, not a correctness check. A wrong but friendly answer can win.
  • Which questions got asked depends on who voted — they may differ from your use.
  • Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.