Tool · curated 24 Sep 2026
GitHub - rudratoshs/buried-injections: 🛡️ Regex catches 0%, Meta's Prompt Guard 2 catches 1% of 629 realistic AgentDojo injection attacks when they're buried in tool output. Reproducible benchmark.
First reported github.com
Coverage timeline
Single-source research — first reported, latest, and curated coincide.
Why it matters
The buried-injections benchmark demonstrates that widely used prompt-injection detectors largely fail when malicious instructions are hidden in tool output, exposing a blind spot defenders relying on these guardrails must account for.
The buried-injections project is a reproducible benchmark that tests prompt-injection detectors against 629 realistic AgentDojo injection attacks embedded ("buried") inside tool output. Results reported show regex detection catching 0% and Meta's Prompt Guard 2 catching roughly 1%, with a 10-detector leaderboard covering ProtectAI DeBERTa, LLM Guard, deepset, fmops, TestSavant, Preamble and Jailbreak-Detector-Large.