VibeSec benchmark · open v1
Can an agent fix a real security bug?
Every task runs a real attack and app test. A model scores only when it blocks the attack without breaking the app.
1,000
security tasks
64.9%
best score
6
models tested
Leaderboard
Secure fixes out of 1,000 verified tasks
1
Claude Opus 4.864.9%649/1000
most common failure: the attack still works
2
Claude Sonnet 4.637.3%373/1000
most common failure: the app breaks
3
Kimi K2.7 Code37.2%372/1000
most common failure: the attack still works
4
GLM 5.233.7%337/1000
most common failure: the attack still works
5
Nemotron 3 Ultra28.9%289/1000
most common failure: the app breaks
6
GPT-OSS 120B11.1%111/1000
most common failure: the app breaks
Want to inspect or run the tasks yourself?
The dataset is public, and the Lab lets you work through an environment.