Research · curated 29 Sep 2026

Defending RAG Against Knowledge Poisoning Using Cross-Encoder Activation Signals

Coverage timeline

29 Jun 2026mlr.press

Single-source research — first reported, latest, and curated coincide.

Why it matters

Knowledge poisoning of a RAG corpus lets adversaries steer LLM outputs toward attacker-chosen targets, and CEG-RAG offers defenders a practical, evaluated method to detect and neutralize such injected passages.

CEG-RAG is a defense framework presented at the 39th Canadian Conference on AI that uses internal activations of a cross-encoder reranker plus multi-instance learning to detect and localize poisoned chunks in Retrieval-Augmented Generation pipelines, then repairs the context before answer generation. Evaluated on MS MARCO, Natural Questions, and HotpotQA, it achieves TPR >85% detection at low false-positive rates and reduces attack success rate by an average of 88.74%, with code and data released.