Research

Understanding and Enhancing Backdoor Persistency in LLM Agent Post-Training

Page published

Publication date unknown · First observed: 11 Oct 2026

Coverage timeline

11 Oct 2026arxiv.orgobservedprimary

Single-source research — one report is available.

Why it matters

PersistBD demonstrates that backdoors inherited from third-party models can survive and even strengthen through benign fine-tuning, a concrete supply-chain risk for teams adapting external models into autonomous coding agents.

Researchers from UIUC, Oxford and others study a supply-chain threat where an attacker ships a backdoored LLM and examine whether the hidden malicious behavior survives a developer's benign post-training (supervised fine-tuning and reinforcement learning) when building software-engineering agents. They find benign SFT erodes backdoors but subsequent RL often preserves or increases attack success, and introduce PersistBD, which refines a backdoored model before release to raise attack success from 20% to 74% after SFT and 76% after SFT-RL on Qwen2.5-Coder-7B while keeping benign performance.