DeepsecBench: evaluating model performance in finding cybersecurity vulnerabilities
Executive Take
Security scanning economics just shifted: cheaper open-weight and mid-tier reasoning models now deliver a large fraction of frontier-model detection quality, so budget allocation should split between continuous cheap sweeps and periodic expensive audits rather than one flat-rate tool.
Executive Summary
Vercel released DeepsecBench, a benchmark scoring AI models on finding vulnerabilities in application code using recall-weighted F2 scoring. GPT-5.6 Sol (xhigh) ranks first at 35.58 for $55.98; Claude Opus 5 (medium) scores 28.36; Kimi K3 (high) scores 17.56 for $12.38. Anthropic's Fable 5 is excluded because it declines security work.
Why It Matters
Technology and cybersecurity leaders need this to design tiered AI-assisted code scanning programs now that model choice, not just budget, determines vulnerability coverage; the OpenAI sandbox breach cited also signals attackers already wield these same capabilities.