Research · curated 6 Oct 2026

Passing the Test You Trained On: Re-evaluating Prompt-Injection Detectors for LLM Agents

Coverage timeline

6 Oct 2026arxiv.orgprimary

Single-source research — first reported, latest, and curated coincide.

Why it matters

Defenders relying on public benchmark scores to choose prompt-injection detectors for LLM agents may be deploying tools that catch only a tiny fraction of real in-agent injections, giving a false sense of security.

The paper 'Passing the Test You Trained On' re-evaluates fifteen prompt-injection detectors (including Meta's Prompt Guard 2) plus two LLM judges by replaying ground-truth tool calls from the AgentDojo and tau-bench agent benchmarks. The authors find detection rankings transfer poorly across benchmarks — the top BIPIA detector catches only 2% of AgentDojo injections at a 1% false-positive rate — and argue evaluations should use the agent's own tool outputs, report detection at low false-positive rates, and audit detector training data.