News · curated 18 Jul 2026

More details on Fable 5’s cyber safeguards and our jailbreak framework

Coverage timeline

18 Jul 2026anthropic.comprimary

Single-source analysis — first reported, latest, and curated coincide.

Why it matters

Anthropic's jailbreak severity framework and classifier design give defenders a shared vocabulary and concrete model of how deployed LLM safeguards decide what dual-use cyber activity to block, which is directly relevant to anyone assessing or hardening AI systems against jailbreaks.

Anthropic's announcement details the cybersecurity safety classifiers shipped with its Claude Fable 5 model — which sort cyber uses into prohibited, high-risk dual-use, low-risk dual-use, and benign categories to block or monitor dangerous requests — and proposes an early-draft AI jailbreak severity framework developed with Glasswing partners, alongside a HackerOne program for researchers to submit cyber jailbreaks. It references related research on Boundary Point Jailbreaking, a black-box attack that evades industry-deployed classifier safeguards.