Analysis
Many-shot jailbreaking
Publication date not yet evaluated · Added · 3 reports anthropic.com
Page published
Coverage timeline
Why it matters
Many-shot jailbreaking turns the long-context windows that ship with every frontier model into an attack surface that evades safety guardrails, and its shared mechanism with in-context learning means naive defenses would break legitimate functionality.
Many-shot jailbreaking (MSJ), researched by Anthropic (Anil, Durmus, Panickssery, Sharma, et al., 2024), exploits large context windows by stuffing hundreds of faux user-assistant turns in which the assistant complies with harmful requests before appending the target query. Attack success follows a power law in shot count — failing at 5 shots but becoming reliable at 256 shots — because MSJ shares the same pattern-extraction mechanism as benign in-context learning, making defenses hard; Anthropic reports a classifier-based prompt-modification defense cutting attack success from 61% to 2%.