Research · curated 9 Aug 2026

When JPEGs Start Giving Orders: A Journey into Multi-modal Prompt Injection

Coverage timeline

21 Jul 2026cybersecuritywriteups.c…

Single-source research — first reported, latest, and curated coincide.

Why it matters

Multi-modal prompt injection via user-supplied images shows that vision-language pipelines can be hijacked through image content, threatening any downstream workflow that trusts AI-generated captions.

A security researcher (Jobson) documents discovering multi-modal prompt injection in an AI-powered application that uses a vision-language model to generate captions from user-supplied images or image URLs. After initial SSRF testing failed, the researcher pursued injecting instructions via image content, whose AI-generated captions feed downstream application workflows.