Research · curated 5 Sep 2026
Modelling and control of jailbreak attacks in AI systems as hybrid cyber–physical security threats | Scientific Reports
First reported nature.com
Coverage timeline
Single-source research — first reported, latest, and curated coincide.
Why it matters
The paper offers defenders a formal control-theoretic framework to measure jailbreak resilience and design adaptive defenses for safety-critical AI systems.
A Scientific Reports paper models AI jailbreak attacks as hybrid cyber-physical security threats using a hybrid automaton framework, treating jailbreaks as coordinated adversarial cyber inputs, physical disruptions, and discrete mode switches. The authors quantify jailbreak success via spurious jump probabilities and safety-barrier deterioration, and propose a detection-and-mitigation scheme based on control barrier functions and MPC-based adaptive defense with Lyapunov stability guarantees.