Research · curated 6 Oct 2026
Passing the Test You Trained On: Re-evaluating Prompt-Injection Detectors for LLM Agents
First reported arxiv.org
Coverage timeline
Single-source research — first reported, latest, and curated coincide.
Why it matters
Defenders relying on public benchmark scores to choose prompt-injection detectors for LLM agents may be deploying tools that catch only a tiny fraction of real in-agent injections, giving a false sense of security.
The paper 'Passing the Test You Trained On' re-evaluates fifteen prompt-injection detectors (including Meta's Prompt Guard 2) plus two LLM judges by replaying ground-truth tool calls from the AgentDojo and tau-bench agent benchmarks. The authors find detection rankings transfer poorly across benchmarks — the top BIPIA detector catches only 2% of AgentDojo injections at a 1% false-positive rate — and argue evaluations should use the agent's own tool outputs, report detection at low false-positive rates, and audit detector training data.