Arena Medicine & Healthcare

In plain words · Which AI people preferred on expert-level medicine and healthcare questions

What does it measure?
Which AI's answers were preferred on expert-level questions in Medicine & Healthcare. Examples: interpreting cases, comparing treatment guidelines, summarising medical papers. These are questions the people who actually do that work would ask, not general ones.
Who evaluates it, and how?
On Arena Intelligence, a user enters a question, two anonymous AIs answer, and the user votes for the better one. Over 82 million votes are aggregated with a statistical model (Bradley-Terry) into an Elo-style score, and this table takes the 'style-controlled' version from the official dataset, which statistically removes the bias toward long, nicely formatted answers. An AI classifier judges each prompt's reasoning depth and expertise and tags only about 5.5% as 'expert' prompts, which are then divided into occupational categories based on the US Bureau of Labor Statistics classification (SOC), launched November 2025. This table counts only prompts classified as 'Medicine & Healthcare'.
Reading the score
This is a head-to-head record, not an exam score (accuracy %). It is computed from the probability of being preferred over other AIs, so there is no maximum, and when a new AI enters, existing scores shift slightly. Typical values are 1,000–1,500; a 100-point gap means the leader is preferred about 64% of the time, a 200-point gap about 76%. Correctness is not counted — a wrong but friendly answer can win — so if accuracy matters, also look at accuracy tests (GPQA, AIME, etc.). Gaps of a few dozen points are effectively the same tier even if the rank differs. This category has few prompts, so top ranks change often; look at score bands rather than ranks.

Last updated: 2026-10-02

Rank

Top score1547Gemini 4 Argon
Models tested96
Last updated2026.10.02
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    Google2026.09.30
    1547
  2. 2
    Anthropic2026.09.22
    1539
  3. 3
    Anthropic2026.09.01
    1523
  4. 4
    Anthropic2026.02.04
    1514
  5. 5
    Anthropic2026.04.16
    1513
  6. 6
    Moonshot AI2026.07.17
    1512
  7. 7
    Meta2026.08.05
    1511
  8. 8
    Google2026.09.03
    1510
  9. 9
    Anthropic2026.07.24
    1506
  10. 10
    Anthropic2026.06.09
    1506

19 of the 37 AIs released in the last 60 days have no score here yet · Claude Haiku 5.5, Decider V1.1 27B, GPT-6 Luna Decisions and 16 more

Source: Arena Intelligence Data sources & removal requests

What to keep in mind

  • It’s a preference vote, not a correctness check. A wrong but friendly answer can win.
  • Which questions got asked depends on who voted — they may differ from your use.
  • Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.