CyberGym
Often citedHigher is better
A UC Berkeley security test: from a description alone, the AI must write code that reproduces 1,507 real software vulnerabilities. It has values measured by the maintainers with open agents such as OpenHands, plus maker-reported values; maker values are with safety mitigations turned off. Scores run from 0 to 100%. Higher is better.
Top score95.1%MiMo-V2.6-Flash
Models tested29
Last updated2026.09.22
Model release date (newest on the right)
Best score so farHigher on the chart is better
RankModelSettingReleasedPrice/1MScore
- 12026.09.22$0.2595.1%
- 22026.09.21$0.7694.0%
- 3max (reasoning_effort=100)2026.09.10$0.7988.1%
- 4High2026.09.20$2.2884.7%

- 5max2026.08.18$3.6084.5%
- 6with tools2026.07.09$1683.6%
- 7with tools · xhigh2026.04.23$2481.8%
- 8Mean Reproduced · with tools · high2026.07.08$580.4%
- 9Mean Reproduced · with tools · high2026.09.21$580.3%
- 10Mean Reproduced · with tools · high2026.08.12$579.7%
The "harness" is the agent program the AI used to carry out the task. The same model can score very differently depending on its harness and reasoning effort.
Vendor-reported numbers use varying setups — treat them as indicative.