Research · curated 15 Jul 2026
DisarmRAG: Stealthy Retriever-Centric Poisoning to Disable Self-Correction in Retrieval-Augmented Generation (Extended Version)
First reported · updated · 2 reports arxiv.org
Coverage timeline
Why it matters
DisarmRAG shows that compromising the retriever can defeat the self-correction behavior defenders rely on in real-world RAG deployments, undermining a key assumed safeguard against knowledge-base poisoning.
DisarmRAG is a research attack framework that poisons the retriever component of Retrieval-Augmented Generation systems—rather than only the knowledge base—to inject anti-self-correction instructions into the LLM context, suppressing models' self-correction ability and forcing attacker-chosen outputs. Using iterative co-optimization and a contrastive-learning-based stealthy model-editing technique, the authors report success rates exceeding 90% across six LLMs and three QA benchmarks while evading detection defenses.