Research · curated 2 Oct 2026

Divide and Inject: Can Agents Reconstruct an Indirect Prompt Injection from Fragments?

Coverage timeline

2 Oct 2026arxiv.orgprimary

Single-source research — first reported, latest, and curated coincide.

Why it matters

AdaLCPI shows agents can reconstruct harmful objectives from scattered benign-looking fragments, defeating defenses that assume a complete malicious instruction must appear in one place and demanding fragment-aware safety evaluation.

The paper introduces AdaLCPI (adaptive long-context prompt injection), an attack that splits a malicious objective into incomplete fragments distributed across an agent's retrieved long-context content and uses a reconstruction cue plus iterative refinement (via OpenEvolve with graded scoring and execution feedback) to make the agent reassemble and execute the hidden instruction. AdaLCPI reaches 61.4% macro-average attack success rate, versus 32.8% for Trojan Hippo-style and 30.0% for AgentVigil baselines.