Research · curated 9 Aug 2026
RoguePrompt: Dual‐Layer Encoding for Self‐Reconstruction to Circumvent LLM Moderation
First reported arxiv.org
Coverage timeline
Single-source research — first reported, latest, and curated coincide.
Why it matters
RoguePrompt demonstrates that layered self-reconstruction encoding reliably circumvents deployed LLM safety controls, giving defenders concrete evidence of where moderation, reconstruction, and execution stages break down.
RoguePrompt is a jailbreak pipeline described in an arXiv paper by researchers at Boston University that partitions a forbidden prompt and applies two nested encodings (Vigenère followed by ROT13) with natural-language reconstruction instructions to evade LLM moderation. Evaluated in a black-box setting against 313 hard-rejected prompts, it achieved 93.93% filter bypass, 79.02% instruction reconstruction, and 70.18% execution, with stage-level measurement of where multistage jailbreaks fail.