Research · curated 21 Jul 2026
PVDetector: Detecting Prompt Injection Attacks on Purpose-Specific LLM Agents through Policy-Violation Concept Analysis
First reported arxiv.org
Coverage timeline
Single-source research — first reported, latest, and curated coincide.
Why it matters
PVDetector offers defenders a detection approach for prompt injection on domain-specific LLM agents by leveraging models' internal awareness of policy violations rather than input-output pattern matching.
PVDetector is a training-free framework presented in an arXiv paper for detecting prompt injection attacks against purpose-specific LLM agents. It works by analyzing the model's hidden activation space for latent policy-violation concepts derived from contrastive pairs of policy-violating and policy-compliant prompts, measuring hidden-state alignment during inference.