Research · curated 31 Jul 2026
Piggybacking on Perception: Stealthy Concurrent Audio Prompt Injections against Multimodal LLM Agents
First reported arxiv.org
Coverage timeline
Single-source research — first reported, latest, and curated coincide.
Why it matters
Audio-driven multimodal agents accept environmental sound beyond user control, so imperceptible piggybacked voice instructions can silently hijack real deployed assistants like Doubao AI Smartphone into malicious actions.
A research paper, "Piggybacking on Perception," demonstrates stealthy concurrent audio prompt injection attacks against multimodal LLM agents, using instruction augmentation and scenario concealment to hide malicious audio instructions inside user speech and hijack agents. The authors build AudioAgentSecurity, a benchmark of 8 scenarios and 10 attack patterns, evaluate 11 agents (including Gemini 3 Pro and GPT-4o-audio) achieving a 69.10% average attack success rate against Gemini 3 Pro, and propose a CADV defense using source separation and cross-modal consistency achieving over 90% detection.