VibeSec benchmark · open v1
Can an agent fix a real security bug?
Every task runs a real attack and app test. A model scores only when it blocks the attack without breaking the app.
1,000
security tasks
55.4%
best score
6
models tested
Leaderboard
Secure fixes out of 1,000 verified tasks
1
Claude Opus 4.855.4%554/1000
most common failure: the attack still works
2
Gemini 3.8 Flash[config]55.1%551/1000
most common failure: the attack still works
3
Claude Sonnet 4.632.8%328/1000
most common failure: the attack still works
4
GLM 5.227.9%279/1000
most common failure: the attack still works
5
Nemotron 3 Ultra27.4%274/1000
most common failure: the attack still works
6
Nemotron 3.5 Lightning25.6%256/1000
most common failure: the attack still works
Want to inspect or run the tasks yourself?
Every task, exploit and model outcome is browsable, and the Lab lets you work through an environment.