Arena Hard Prompts

In plain words · Which AI people preferred on the hardest questions

What does it measure?
The ranking recomputed on hard questions only, with easy ones removed. It shows which AI was preferred on questions needing expertise, complex reasoning and accuracy.
Who evaluates it, and how?
On Arena Intelligence, a user enters a question, two anonymous AIs answer, and the user votes for the better one. Over 82 million votes are aggregated with a statistical model (Bradley-Terry) into an Elo-style score, and this table takes the 'style-controlled' version from the official dataset, which statistically removes the bias toward long, nicely formatted answers. Prompts are rated on seven criteria — specificity, domain knowledge, reasoning steps, technical accuracy and more — and only the top third or so are counted.
Reading the score
This is a head-to-head record, not an exam score (accuracy %). It is computed from the probability of being preferred over other AIs, so there is no maximum, and when a new AI enters, existing scores shift slightly. Typical values are 1,000–1,500; a 100-point gap means the leader is preferred about 64% of the time, a 200-point gap about 76%. Correctness is not counted — a wrong but friendly answer can win — so if accuracy matters, also look at accuracy tests (GPQA, AIME, etc.). Gaps of a few dozen points are effectively the same tier even if the rank differs. An AI whose score differs a lot from the overall ranking performs differently on easy versus hard questions.

Last updated: 2026-10-02

Rank

Top score1551Gemini 4 Argon
Models tested97
Last updated2026.10.02
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    Google2026.09.30
    1551
  2. 2
    Anthropic2026.09.22
    1534
  3. 3
    Anthropic2026.06.09
    1530
  4. 4
    Anthropic2026.02.04
    1526
  5. 5
    Anthropic2026.09.01
    1520
  6. 6
    Anthropic2026.04.16
    1520
  7. 7
    Moonshot AI2026.07.17
    1517
  8. 8
    Meta2026.09.02
    1517
  9. 9
    Google2026.09.03
    1516
  10. 10
    Anthropic2026.07.24
    1514

18 of the 37 AIs released in the last 60 days have no score here yet · Claude Haiku 5.5, Decider V1.1 27B, GPT-6 Luna Decisions and 15 more

Source: Arena Intelligence Data sources & removal requests

What to keep in mind

  • It’s a preference vote, not a correctness check. A wrong but friendly answer can win.
  • Which questions got asked depends on who voted — they may differ from your use.
  • Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.