Research · curated 9 Aug 2026

RoguePrompt: Dual‐Layer Encoding for Self‐Reconstruction to Circumvent LLM Moderation

Coverage timeline

9 Aug 2026arxiv.orgprimary

Single-source research — first reported, latest, and curated coincide.

Why it matters

RoguePrompt demonstrates that layered self-reconstruction encoding reliably circumvents deployed LLM safety controls, giving defenders concrete evidence of where moderation, reconstruction, and execution stages break down.

RoguePrompt is a jailbreak pipeline described in an arXiv paper by researchers at Boston University that partitions a forbidden prompt and applies two nested encodings (Vigenère followed by ROT13) with natural-language reconstruction instructions to evade LLM moderation. Evaluated in a black-box setting against 313 hard-rejected prompts, it achieved 93.93% filter bypass, 79.02% instruction reconstruction, and 70.18% execution, with stage-level measurement of where multistage jailbreaks fail.