Research · curated 23 Jul 2026
Dynamic Defense Profiling Enables Cognitive Jailbreak of Text-to-Image Models
First reported arxiv.org
Coverage timeline
Single-source research — first reported, latest, and curated coincide.
Why it matters
MIND demonstrates that adaptive, feedback-driven jailbreaks can defeat layered safety filters on both open-source and commercial text-to-image models at very high success rates, undermining current NSFW guardrails.
Researchers present MIND, a cognitive jailbreak framework that models a text-to-image system's latent defense mechanisms as a belief-state inference problem, interpreting multi-modal feedback (textual refusal, visual blocking, semantic sanitization) to iteratively craft adversarial prompts that produce NSFW content. Using a Multi-modal Judge, Defense Profiler, and Meta-Memory module, MIND reports a 95.62% attack success rate against defended Stable Diffusion v1.5 and up to 91.58% against commercial T2I systems like Wan-2.5.