Research · curated 29 Jul 2026
MJ: Multi-turn LLM Jailbreaking via Decomposed Credit Assignment
First reported arxiv.org
Coverage timeline
Single-source research — first reported, latest, and curated coincide.
Why it matters
Multi-turn jailbreak learning like DC-GRPO shows how automated attackers can reliably bypass LLM safety alignment across realistic conversational settings, raising the bar for red-teaming and defensive alignment.
The paper "MJ: Multi-turn LLM Jailbreaking via Decomposed Credit Assignment" introduces DC-GRPO, a turn-level credit assignment framework for training reinforcement-learning attackers that jailbreak LLMs across multi-turn conversations. The authors report attack success rates (ASR5@3) of roughly 98% across multiple victim LLMs and benchmarks, substantially outperforming prior methods such as SEMA and TROJail.