We burned 11.7 billion tokens to benchmark the cyber capabilities of 10 AI models with three attempts each, given 32 fresh off-the-shelf vulns to rediscover.
We burned 11.7bn tokens to find the best cyber AI model
We burned 11.7 billion tokens to benchmark the cyber capabilities of 10 AI models with three attempts each, given 32 fresh off-the-shelf vulns to rediscover.
Aikido Security
Publisher
Aug 21, 2026 at 9:15 AM UTC · Updated 11 giờ trước · 9 phút đọc

Market Impact
SOL+7.24%$95.55
Last Updated
11 giờ trước
This benchmark is an evolution of our earlier known-CVE benchmark with a fresh, harder dataset, more models, and a closer look at what they find, how reliably they find it, and their quirks and trade-offs.
With this, we add GLM-5.3, DeepSeek V4 Pro 0813, DeepSeek V4 Flash 0731, Qwen3.8-Max, Kimi K3, and Grok 4.6 to the lineup.
TL;DR:
- DeepSeek V4 Pro 0813 finds the most vulnerabilities. Pooling three runs reaches 28 of 32 vulnerabilities.
- The most expensive model is not required. Three DeepSeek Pro runs cost about $295 and outperform Opus 5, Grok 4.6, or Sol. Three Flash runs cost $108 and reach 24 matching Grok’s best individual pass for less than a quarter of the cost.
- Models are inconsistent at recall; repetition remediates it. Single runs miss the breadth of findings, but pooling across runs fills this gap. DeepSeek Pro finds 17 vulnerabilities on its first pass but 28 across three.
- Open-source models now outperform the public frontier. DeepSeek V4 Pro topped every public closed model we tested on pooled vulnerability recall. Qwen, Kimi, and GLM-5.3 followed with strong consistency of findings without losing recall. Harnessed correctly, open models can now compete directly with the closed frontiers.
- The price of cheap coverage is noise. The open models caught up with the frontier at much cheaper rates but also produced the most false leads for the pipeline to reject.
Market Context
Article Intelligence
Sponsored
AdNewsLayer Premium
Unlock deeper intelligence.
Ad-free reading, exclusive research, and real-time onchain insights.
Go Premium
