Research · curated 18 Jul 2026
Rethinking Jailbreak Detection of Large Vision Language ...
First reported aclanthology.org
Coverage timeline
Single-source research — first reported, latest, and curated coincide.
Why it matters
RCS offers defenders a lightweight, generalizable method to detect adversarial-image and cross-modal prompt-injection jailbreaks against vision-language models before they bypass safety mechanisms in deployed systems.
Researchers from Washington University in St. Louis and Texas A&M propose Representational Contrastive Scoring (RCS), a framework for detecting multimodal jailbreak attacks against Large Vision-Language Models by inspecting the models' internal representations. The two instantiations, MCD and KCD, claim state-of-the-art detection performance and better generalization to unseen attack types while reducing over-rejection, with code released on GitHub.