Research · curated 15 Jul 2026

DisarmRAG: Stealthy Retriever-Centric Poisoning to Disable Self-Correction in Retrieval-Augmented Generation (Extended Version)

Coverage timeline

15 Jul 2026arxiv.org 26 Jul 2026arxiv.orgprimary

Why it matters

DisarmRAG shows that compromising the retriever can defeat the self-correction behavior defenders rely on in real-world RAG deployments, undermining a key assumed safeguard against knowledge-base poisoning.

DisarmRAG is a research attack framework that poisons the retriever component of Retrieval-Augmented Generation systems—rather than only the knowledge base—to inject anti-self-correction instructions into the LLM context, suppressing models' self-correction ability and forcing attacker-chosen outputs. Using iterative co-optimization and a contrastive-learning-based stealthy model-editing technique, the authors report success rates exceeding 90% across six LLMs and three QA benchmarks while evading detection defenses.