Research · curated 18 Jul 2026

Rethinking Jailbreak Detection of Large Vision Language ...

Coverage timeline

18 Jul 2026aclanthology.org

Single-source research — first reported, latest, and curated coincide.

Why it matters

RCS offers defenders a lightweight, generalizable method to detect adversarial-image and cross-modal prompt-injection jailbreaks against vision-language models before they bypass safety mechanisms in deployed systems.

Researchers from Washington University in St. Louis and Texas A&M propose Representational Contrastive Scoring (RCS), a framework for detecting multimodal jailbreak attacks against Large Vision-Language Models by inspecting the models' internal representations. The two instantiations, MCD and KCD, claim state-of-the-art detection performance and better generalization to unseen attack types while reducing over-rejection, with code released on GitHub.