Tool · curated 3 Oct 2026
ZeroLeaks/shield-small
First reported huggingface.co
Coverage timeline
Single-source research — first reported, latest, and curated coincide.
Why it matters
Shield Small gives defenders a locally runnable, downloadable guardrail to screen untrusted content feeding LLM agents, helping mitigate indirect prompt injection in RAG and tool-calling pipelines.
Shield Small is a 118M-parameter ONNX classifier published on Hugging Face by ZeroLeaks for detecting prompt injection and jailbreak attempts in untrusted text before an AI agent reads it (retrieved documents, tool results, web pages, tool descriptions). It runs locally on CPU, returns an injection score and binary verdict, and reports 84.3% mean category-balanced accuracy and 71.8% attack recall when combined with rules v2.