Research · curated 23 Jul 2026

Dynamic Defense Profiling Enables Cognitive Jailbreak of Text-to-Image Models

Coverage timeline

23 Jul 2026arxiv.orgprimary

Single-source research — first reported, latest, and curated coincide.

Why it matters

MIND demonstrates that adaptive, feedback-driven jailbreaks can defeat layered safety filters on both open-source and commercial text-to-image models at very high success rates, undermining current NSFW guardrails.

Researchers present MIND, a cognitive jailbreak framework that models a text-to-image system's latent defense mechanisms as a belief-state inference problem, interpreting multi-modal feedback (textual refusal, visual blocking, semantic sanitization) to iteratively craft adversarial prompts that produce NSFW content. Using a Multi-modal Judge, Defense Profiler, and Meta-Memory module, MIND reports a 95.62% attack success rate against defended Stable Diffusion v1.5 and up to 91.58% against commercial T2I systems like Wan-2.5.