Research · curated 12 Sep 2026
AgentHijack: Visual Patch Attacks on Multimodal Computer-Use Agents
First reported arxiv.org
Coverage timeline
Single-source research — first reported, latest, and curated coincide.
Why it matters
AgentHijack shows adversarial visual patches can propagate through the full screenshot-to-execution pipeline of computer-use agents to trigger real environmental actions like terminal commands, expanding the attack surface for anyone deploying multimodal GUI agents.
AgentHijack is a research framework evaluating image-triggered command injection against multimodal computer-use agents (CUAs), where optimized local visual patches embedded in web pages induce malicious terminal commands. Across 600 online cases spanning five open-source GUI-agent/VLM backends, the authors report T-ASR of 84.5%, TAPR of 47.0%, and end-to-end attack success (E2E-ASR) of 20.3%, with some agents executing a malicious command before continuing the benign task.