Research

Covert Visual Prompt Injection against Commercial Multimodal Large Language Models

Page published

Coverage timeline

discovered arxiv.org primary 26 Aug 2026deeplearn.orgobserved

Single-source research — one report is available.

Why it matters

Covert visual prompt injection lets attackers hide malicious instructions inside seemingly benign images that fool deployed multimodal AI systems without any human-visible cues, expanding the prompt-injection attack surface for defenders.

A research paper by Meiwen Ding and colleagues presents a covert visual prompt injection attack against commercial closed-source multimodal large language models (MLLMs). The method embeds imperceptible adversarial perturbations and a bounded text overlay into an input image, iteratively optimizing feature alignment with malicious visual and textual targets to smuggle instructions past human observers and transfer across multiple MLLMs.