Research · curated 29 Sep 2026
Defending RAG Against Knowledge Poisoning Using Cross-Encoder Activation Signals
First reported mlr.press
Coverage timeline
Single-source research — first reported, latest, and curated coincide.
Why it matters
Knowledge poisoning of a RAG corpus lets adversaries steer LLM outputs toward attacker-chosen targets, and CEG-RAG offers defenders a practical, evaluated method to detect and neutralize such injected passages.
CEG-RAG is a defense framework presented at the 39th Canadian Conference on AI that uses internal activations of a cross-encoder reranker plus multi-instance learning to detect and localize poisoned chunks in Retrieval-Augmented Generation pipelines, then repairs the context before answer generation. Evaluated on MS MARCO, Natural Questions, and HotpotQA, it achieves TPR >85% detection at low false-positive rates and reduces attack success rate by an average of 88.74%, with code and data released.