Analysis · curated 20 Aug 2026
LLM Jailbreak: How Prompts Bypass Guardrails
First reported casrai.org
Coverage timeline
Single-source analysis — first reported, latest, and curated coincide.
Why it matters
The CASRAI jailbreak glossary entry gives defenders a shared taxonomy and vocabulary for classifying LLM safety-bypass techniques, but it is evergreen reference material rather than a new threat or finding.
CASRAI's dictionary entry "Jailbreak (LLM)" defines a jailbreak as a prompt or interaction pattern that causes a language model to bypass its safety training and produce refused outputs, cataloguing techniques such as role-play framings, multi-turn manipulation, encoding tricks (base64, ROT13), and adversarial-suffix attacks. The reference material distinguishes jailbreaks from prompt injection and describes mitigations like RLHF, constitutional AI, and red-team evaluation.