Research · curated 25 Jul 2026
Local AI for Penetration Testing & Research
First reported projectblack.io
Coverage timeline
Single-source research — first reported, latest, and curated coincide.
Why it matters
Benchmarking agentic and local-AI vulnerability-discovery workflows shows defenders how effective (and inefficient) autonomous AI pentesting tools currently are and where custom harnesses outperform off-the-shelf agents.
A researcher benchmarked four approaches for discovering a known authenticated LFI vulnerability (CVE-2026-12194 in PHPIPAM): Semgrep, the agentic Strix pentest agent with GLM 5.1, a cloud SOTA model with a code-review skill, and a local AI model with a custom harness. Only the local AI + custom harness reliably found the bug, with the author concluding that the harness/approach matters more than the underlying model, while the Strix agent consumed ~60 million tokens over ~12 hours without success.