Tool · curated 3 Oct 2026

ZeroLeaks/shield-small

Coverage timeline

3 Oct 2026huggingface.co

Single-source research — first reported, latest, and curated coincide.

Why it matters

Shield Small gives defenders a locally runnable, downloadable guardrail to screen untrusted content feeding LLM agents, helping mitigate indirect prompt injection in RAG and tool-calling pipelines.

Shield Small is a 118M-parameter ONNX classifier published on Hugging Face by ZeroLeaks for detecting prompt injection and jailbreak attempts in untrusted text before an AI agent reads it (retrieved documents, tool results, web pages, tool descriptions). It runs locally on CPU, returns an injection score and binary verdict, and reports 84.3% mean category-balanced accuracy and 71.8% attack recall when combined with rules v2.