Research · curated 5 Oct 2026
CounterSteer: Suppressing Indirect Prompt Injection with Activation Steering
First reported arxiv.org
Coverage timeline
Single-source research — first reported, latest, and curated coincide.
Why it matters
CounterSteer offers defenders a detection-free, always-on activation-steering mechanism to blunt indirect prompt injection in LLM agents without fine-tuning served weights or sacrificing much benign utility.
CounterSteer, presented by Mark Russinovich in an arXiv paper, is an inference-time defense against indirect prompt injection that subtracts a fitted residual-stream direction from every tool-result token during prefill, requiring no fine-tuning or auxiliary model. Across five open-weights models (8B-106B), it cuts held-out attack success from 0.21-1.00 to 0.00-0.17 and AgentDojo compromise from 0.10-0.49 to 0.006-0.079 while retaining 93-100% benign utility, though parameter-manipulation attacks are only partially resisted.