CyberGym

Often citedHigher is better

A UC Berkeley security test: from a description alone, the AI must write code that reproduces 1,507 real software vulnerabilities. It has values measured by the maintainers with open agents such as OpenHands, plus maker-reported values; maker values are with safety mitigations turned off. Scores run from 0 to 100%. Higher is better.

Top score95.1%MiMo-V2.6-Flash
Models tested29
Last updated2026.09.22
Model release date (newest on the right)
Best score so farHigher on the chart is better
  1. 195.1%
  2. 294.0%
  3. 3
    DeepSeek2026.09.10 · max (reasoning_effort=100)
    88.1%
  4. 4
    Stepfun2026.09.20 · High
    84.7%
  5. 5
    Z.ai2026.08.18 · max
    84.5%
  6. 6
    OpenAI2026.07.09 · with tools
    83.6%
  7. 7
    OpenAI2026.04.23 · with tools · xhigh
    81.8%
  8. 8
    xAI2026.07.08 · Mean Reproduced · with tools · high
    80.4%
  9. 9
    xAI2026.09.21 · Mean Reproduced · with tools · high
    80.3%
  10. 10
    xAI2026.08.12 · Mean Reproduced · with tools · high
    79.7%

The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.

Vendor-reported numbers use varying setups — treat them as indicative.