Research

Jailbreaking in the Haystack

Page published

Coverage timeline

25 Aug 2026arxiv.orgobservedprimary

Single-source research — one report is available.

Why it matters

Ninja reveals that expanded context windows in modern long-context and agentic LLMs are themselves a fundamental safety vulnerability, enabling low-resource, transferable, and hard-to-detect jailbreaks that defenders must account for.

Researchers at Carnegie Mellon introduce Ninja (Needle-in-haystack jailbreak attack), a method that jailbreaks aligned long-context language models by embedding harmful goals within long, benign, model-generated context and carefully positioning the goal. Experiments on the HarmBench safety benchmark show Ninja significantly raises attack success rates against LLaMA, Qwen, Mistral, and Gemini, and demonstrate that increasing context length is compute-optimal versus best-of-N jailbreaking. Code is published at github.com/AR-FORUM/NINJA_Attack.