Research · curated 21 Jul 2026

PVDetector: Detecting Prompt Injection Attacks on Purpose-Specific LLM Agents through Policy-Violation Concept Analysis

Coverage timeline

15 Jul 2026arxiv.orgprimary

Single-source research — first reported, latest, and curated coincide.

Why it matters

PVDetector offers defenders a detection approach for prompt injection on domain-specific LLM agents by leveraging models' internal awareness of policy violations rather than input-output pattern matching.

PVDetector is a training-free framework presented in an arXiv paper for detecting prompt injection attacks against purpose-specific LLM agents. It works by analyzing the model's hidden activation space for latent policy-violation concepts derived from contrastive pairs of policy-violating and policy-compliant prompts, measuring hidden-state alignment during inference.