Research
Understanding and Enhancing Backdoor Persistency in LLM Agent Post-Training
Publication date unknown · Discovered arxiv.org
Page published
Publication date unknown · First observed: 11 Oct 2026
Coverage timeline
Single-source research — one report is available.
Why it matters
PersistBD demonstrates that backdoors inherited from third-party models can survive and even strengthen through benign fine-tuning, a concrete supply-chain risk for teams adapting external models into autonomous coding agents.
Researchers from UIUC, Oxford and others study a supply-chain threat where an attacker ships a backdoored LLM and examine whether the hidden malicious behavior survives a developer's benign post-training (supervised fine-tuning and reinforcement learning) when building software-engineering agents. They find benign SFT erodes backdoors but subsequent RL often preserves or increases attack success, and introduce PersistBD, which refines a backdoored model before release to raise attack success from 20% to 74% after SFT and 76% after SFT-RL on Qwen2.5-Coder-7B while keeping benign performance.