Analysis · curated 20 Aug 2026

LLM Jailbreak: How Prompts Bypass Guardrails

Coverage timeline

20 Aug 2026casrai.org

Single-source analysis — first reported, latest, and curated coincide.

Why it matters

The CASRAI jailbreak glossary entry gives defenders a shared taxonomy and vocabulary for classifying LLM safety-bypass techniques, but it is evergreen reference material rather than a new threat or finding.

CASRAI's dictionary entry "Jailbreak (LLM)" defines a jailbreak as a prompt or interaction pattern that causes a language model to bypass its safety training and produce refused outputs, cataloguing techniques such as role-play framings, multi-turn manipulation, encoding tricks (base64, ROT13), and adversarial-suffix attacks. The reference material distinguishes jailbreaks from prompt injection and describes mitigations like RLHF, constitutional AI, and red-team evaluation.