Research · curated 11 Sep 2026
An Experimental Evaluation of Multimodal PromptInjection Attacks on Agentic AI Frameworks
First reported arxiv.org
Coverage timeline
Single-source research — first reported, latest, and curated coincide.
Why it matters
Multimodal prompt injection via images and audio lets attackers slip instructions into agent context without going through the user, and this study shows perceptual channels beyond vision are narrower but much less defended in tool-calling agents that can delete files or send email.
MMPIBench is a reproducible benchmark presented in an experimental evaluation of multimodal prompt injection attacks against six agentic AI frameworks (LangGraph, CrewAI, AutoGen, OpenAI Agents SDK, Semantic Kernel, LlamaIndex). Across 720 runs spanning six visual carriers (OCR text, overlays, EXIF metadata, QR codes, fake interfaces, hybrids), attacks completed in ~1% of runs but were attempted in 12.8%, with the model mattering more than the framework; extending to audio, attacks completed in 49% of applicable cells and 75% for one model.