Research · curated 15 Aug 2026
When Agents Go Rogue: Activation-Based Detection of Malicious Behaviors in Multi-Agent Systems
First reported arxiv.org
Coverage timeline
Single-source research — first reported, latest, and curated coincide.
Why it matters
AcMAS addresses a growing defensive gap in LLM multi-agent systems, where a single compromised or influenced agent can subvert an entire collaborative system through semantically stealthy attacks that evade conventional graph-based defenses.
Researchers at Worcester Polytechnic Institute presented AcMAS, an activation-based framework for detecting stealthy malicious behaviors in LLM-based multi-agent systems (MAS), at ICML 2026 (arXiv:2607.06807). AcMAS analyzes internal reasoning states in the activation space of local agents to detect compromised agents without relying on explicit interaction graphs, reporting large F1 improvements over graph-based baselines in both synchronous (0.94 vs 0.72) and asynchronous (0.93 vs 0.38) settings, and can help restore compromised agents rather than isolating them.