Research · curated 4 Aug 2026
State-Dependent Safety Failures in Multi-Turn Language Model Interaction
First reported arxiv.org
Coverage timeline
Single-source research — first reported, latest, and curated coincide.
Why it matters
STAR demonstrates that single-turn safety evaluations miss a whole class of multi-turn jailbreak failures, meaning deployed LLMs defenders consider safe can be steered into harmful output through conversational state evolution.
A research paper introduces STAR, a state-oriented diagnostic framework that treats multi-turn dialogue history as a state transition operator to analyze how safety-aligned language models drift toward safety failure across conversational trajectories. Across multiple frontier LLMs, the authors show that models robust under single-turn evaluation can undergo rapid, reproducible safety collapse under structured multi-turn interaction, driven by monotonic drift away from refusal representations and role-conditioned phase transitions.