ExploitBench

Often citedHigher is better

A Carnegie Mellon security test: the agent tries to exploit 41 real, since-patched vulnerabilities in Chrome's V8 JavaScript engine. Progress is split into 16 goals, from reaching the bug to running arbitrary code, and the score is the average share of goals reached. Scores run from 0 to 100%. Higher is better.

Top score47.4%GPT-5.5
Models tested8
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 1
    OpenAI2026.04.23
    47.4%
  2. 2
    Anthropic2026.04.16
    26.5%
  3. 3
    Google2026.02.19
    26.1%
  4. 4
    Anthropic2026.02.17
    23.6%
  5. 5
    Moonshot AI2026.04.20
    18.4%
  6. 6
    Z.ai2026.04.07
    18.1%
  7. 7
    Anthropic2025.10.15
    13.7%
  8. 8
    MiniMax2026.03.18
    13.3%