Research · curated 29 Jul 2026

ALIBI: Adaptive Agentic Attacks on LLM-Based Vulnerability Detectors via Adversarial Code Comments

Coverage timeline

29 Jul 2026arxiv.orgprimary

Single-source research — first reported, latest, and curated coincide.

Why it matters

ALIBI shows that LLM-based vulnerability detectors and code-review agents can be systematically fooled by adversarial comments, undermining trust in AI-driven security tooling and highlighting the need for architectural isolation and comment sanitization.

ALIBI is an automated adaptive black-box attack framework that evades LLM-based vulnerability detectors by inserting adversarial source-code comments that steer detector reasoning or fabricate external tool results without changing program behavior. Evaluated against four detectors including frontier multi-agent systems, it achieves attack success rates exceeding 90% across 125 real-world null-pointer dereference vulnerabilities, reaching 100% on one system, while prompt-level defenses offer limited robustness.