Research · curated 15 Aug 2026

When Agents Go Rogue: Activation-Based Detection of Malicious Behaviors in Multi-Agent Systems

Coverage timeline

discovered arxiv.org primary 10 Aug 2026techxplore.com

Single-source research — first reported, latest, and curated coincide.

Why it matters

AcMAS addresses a growing defensive gap in LLM multi-agent systems, where a single compromised or influenced agent can subvert an entire collaborative system through semantically stealthy attacks that evade conventional graph-based defenses.

Researchers at Worcester Polytechnic Institute presented AcMAS, an activation-based framework for detecting stealthy malicious behaviors in LLM-based multi-agent systems (MAS), at ICML 2026 (arXiv:2607.06807). AcMAS analyzes internal reasoning states in the activation space of local agents to detect compromised agents without relying on explicit interaction graphs, reporting large F1 improvements over graph-based baselines in both synchronous (0.94 vs 0.72) and asynchronous (0.93 vs 0.38) settings, and can help restore compromised agents rather than isolating them.