Arena Hard Prompts
In plain words · Which AI people preferred on the hardest questions
- What does it measure?
- The ranking recomputed on hard questions only, with easy ones removed. It shows which AI was preferred on questions needing expertise, complex reasoning and accuracy.
- Who evaluates it, and how?
- On Arena Intelligence, a user enters a question, two anonymous AIs answer, and the user votes for the better one. Over 82 million votes are aggregated with a statistical model (Bradley-Terry) into an Elo-style score, and this table takes the 'style-controlled' version from the official dataset, which statistically removes the bias toward long, nicely formatted answers. Prompts are rated on seven criteria — specificity, domain knowledge, reasoning steps, technical accuracy and more — and only the top third or so are counted.
- Reading the score
- This is a head-to-head record, not an exam score (accuracy %). It is computed from the probability of being preferred over other AIs, so there is no maximum, and when a new AI enters, existing scores shift slightly. Typical values are 1,000–1,500; a 100-point gap means the leader is preferred about 64% of the time, a 200-point gap about 76%. Correctness is not counted — a wrong but friendly answer can win — so if accuracy matters, also look at accuracy tests (GPQA, AIME, etc.). Gaps of a few dozen points are effectively the same tier even if the rank differs. An AI whose score differs a lot from the overall ranking performs differently on easy versus hard questions.
Last updated: 2026-10-02
Rank
Top score1551Gemini 4 Argon
Models tested97
Last updated2026.10.02
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelReleasedPrice/1MScore
- 12026.09.30$81551

- 22026.09.22$161534
- 32026.06.09$401530
- 42026.02.04$201526
- 52026.09.01$401520
- 62026.04.16$201520
- 72026.07.17$111517
- 82026.09.02$3.501517
- 92026.09.03$31516

- 102026.07.24$201514
18 of the 37 AIs released in the last 60 days have no score here yet · Claude Haiku 5.5, Decider V1.1 27B, GPT-6 Luna Decisions and 15 more
Source: Arena Intelligence Data sources & removal requestsWhat to keep in mind
- It’s a preference vote, not a correctness check. A wrong but friendly answer can win.
- Which questions got asked depends on who voted — they may differ from your use.
- Scores labelled 'Self-reported' were published by the AI maker itself, not measured by a third party.