Analysis · curated 26 Aug 2026
Many-shot jailbreaking
First reported anthropic.com
Coverage timeline
Single-source analysis — first reported, latest, and curated coincide.
Why it matters
Many-shot jailbreaking demonstrates that every context-window increase ships a larger jailbreak surface, and because it abuses in-context learning itself the technique cannot simply be patched away, making it a durable risk for anyone deploying long-context LLMs.
An explainer on many-shot jailbreaking synthesizes Anthropic's 2024 research (Anil et al., NeurIPS 2024) showing that padding a prompt with many fabricated assistant-complies-with-harmful-requests exchanges before a real query erodes safety alignment, with effectiveness scaling as a power law in the number of demonstrations. The piece frames the technique as exploiting in-context learning and growing context windows rather than a patchable bug, and notes it worked across Claude 2.0, GPT-3.5/4, Llama 2 70B, and Mistral 7B.