VibeSec benchmark · open v1

Can an agent fix a real security bug?

Every task runs a real attack and app test. A model scores only when it blocks the attack without breaking the app.

1,000

security tasks

64.9%

best score

6

models tested

Leaderboard

Secure fixes out of 1,000 verified tasks

live results
1
Claude Opus 4.864.9%649/1000

most common failure: the attack still works

2
Claude Sonnet 4.637.3%373/1000

most common failure: the app breaks

3
Kimi K2.7 Code37.2%372/1000

most common failure: the attack still works

4
GLM 5.233.7%337/1000

most common failure: the attack still works

5
Nemotron 3 Ultra28.9%289/1000

most common failure: the app breaks

6
GPT-OSS 120B11.1%111/1000

most common failure: the app breaks

Want to inspect or run the tasks yourself?

The dataset is public, and the Lab lets you work through an environment.