VibeSec benchmark · open v1

Can an agent fix a real security bug?

Every task runs a real attack and app test. A model scores only when it blocks the attack without breaking the app.

1,000

security tasks

55.4%

best score

6

models tested

Leaderboard

Secure fixes out of 1,000 verified tasks

live results
1
Claude Opus 4.855.4%554/1000

most common failure: the attack still works

2
Gemini 3.8 Flash[config]55.1%551/1000

most common failure: the attack still works

3
Claude Sonnet 4.632.8%328/1000

most common failure: the attack still works

4
GLM 5.227.9%279/1000

most common failure: the attack still works

5
Nemotron 3 Ultra27.4%274/1000

most common failure: the attack still works

6
Nemotron 3.5 Lightning25.6%256/1000

most common failure: the attack still works

Want to inspect or run the tasks yourself?

Every task, exploit and model outcome is browsable, and the Lab lets you work through an environment.