HealthBench
Often citedHigher is better
OpenAI's medical-conversation test: responses to 5,000 multi-turn conversations with patients and clinicians are graded against rubrics written by physicians. Values here are standardized on the length-adjusted score. Scores run from 0 to 100%. Higher is better.
This ranking is built only from scores the model makers published themselves. Test settings such as tool use and reasoning effort differ by maker, so check each score's setting before comparing models directly.
Top score60.6%Claude Opus 5.5
Models tested12
Last updated2026.09.22
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelSettingReleasedPrice/1MScore
- 1length-adjusted · no tools · max (adaptive thinking)2026.09.22$1660.6%Anthropic2026.09.22 · length-adjusted · no tools · max (adaptive thinking)
- 2length-adjusted · no tools · max (adaptive thinking)2026.09.01$4060.0%Anthropic2026.09.01 · length-adjusted · no tools · max (adaptive thinking)
- 3length-adjusted · no tools · max (adaptive thinking)2026.05.27$2059.3%Anthropic2026.05.27 · length-adjusted · no tools · max (adaptive thinking)
- 4length-adjusted2026.09.04$4058.3%
- 5length-adjusted · no tools · max (adaptive thinking)2026.07.24$2057.8%Anthropic2026.07.24 · length-adjusted · no tools · max (adaptive thinking)
- 6length-adjusted2026.07.09$9.5057.0%
- 6length-adjusted2026.07.09$1657.0%
- 8length-adjusted2026.04.23$2456.5%
- 9length-adjusted2026.07.09$0.9555.8%
- 10length-adjusted2026.09.22$0.4054.5%
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.
Vendor-reported numbers use varying setups — treat them as indicative.