LinkedIn·Friday, 21 August 2026·5d ago
👀🔥 We burned 11.7 billion tokens to find out which AI models can actually find vulnerabilities. 10 models, three runs each, 32 freshly…
Aikido Security
41,357 followers
👀🔥 We burned 11.7 billion tokens to find out which AI models can actually find vulnerabilities. 10 models, three runs each, 32 freshly disclosed CVEs to rediscover from source.
Open weights won this round. DeepSeek V4 Pro found the most, 28 of 32 pooled across three runs, ahead of Opus 5, Grok 4.6, and GPT-5.6 Sol. Qwen, Kimi, and GLM-5.3 followed close behind, holding recall without losing consistency.
Models are inconsistent at recall, and repetition fixes it. DeepSeek Pro found just 17 on its first run but 28 across three, because each run explores different paths and you keep everything they find.
And it wasn't the priciest. Three DeepSeek Pro runs cost about $295 and beat any single pass of Opus, Grok, or Sol at $450 to $590. Even DeepSeek Flash, at $108 for three runs, reached 24, matching Grok's best single pass for under a quarter of the cost.
The catch is noise. The open models reached the frontier at a fraction of the cost, but produced the most false leads to clear. So the real question isn't which model is best. It's where each one fits, and none of it works without the harness around the model.
Full breakdown of all 10, ranked with costs and trade-offs, by Debarshi and Philippe: https://lnkd.in/dMu2Z7EF
♥ 238💬 15↻ 17
View on LinkedIn Cross-referenced
Related on the wire
npm shipped Trusted Publishing in late 2025. It's free and takes about ten minutes to set up (not to mention, it also blocks an entire…
🌶️🌶️🌶️
Introducing Android Pentests 🚀 Autonomous AI agents log into your app and test it the way an attacker would. One assessment covers the…
We're #hiring a new AI Engineer (Infrastructure Pentest) in Ghent, Flemish Region. Apply today or share this post with your network.
Aikido Security achieves ISO 42001:2023 certification, the international standard for AI governance. 🌟 Independently audited AI governance…
Earlier this year, someone’s OpenClaw agent reportedly hacked a gym's booking system in Australia, just from being asked to help book a…