Research · curated 25 Jul 2026

Local AI for Penetration Testing & Research

Coverage timeline

28 Jun 2026projectblack.io

Single-source research — first reported, latest, and curated coincide.

Why it matters

Benchmarking agentic and local-AI vulnerability-discovery workflows shows defenders how effective (and inefficient) autonomous AI pentesting tools currently are and where custom harnesses outperform off-the-shelf agents.

A researcher benchmarked four approaches for discovering a known authenticated LFI vulnerability (CVE-2026-12194 in PHPIPAM): Semgrep, the agentic Strix pentest agent with GLM 5.1, a cloud SOTA model with a code-review skill, and a local AI model with a custom harness. Only the local AI + custom harness reliably found the bug, with the author concluding that the harness/approach matters more than the underlying model, while the Strix agent consumed ~60 million tokens over ~12 hours without success.