Tool · curated 26 Aug 2026
patronus-studio/wolf-defender-prompt-injection
First reported huggingface.co
Coverage timeline
Single-source analysis — first reported, latest, and curated coincide.
Why it matters
Wolf Defender gives defenders a downloadable, deployable guardrail model to screen untrusted content for prompt injection and jailbreak attempts before it reaches tool-using LLM agents.
Wolf Defender is a multilingual ModernBERT-based (mmBERT-base) binary classifier published on Hugging Face by Patronus that detects prompt injections and jailbreak-style instructions before untrusted content reaches an LLM. The v2 release provides a 2,048-token context window, ONNX deployment variants, and benchmark results showing improved specificity on hard-benign inputs, and is intended as a local guardrail layer for AI agents, chatbots, and retrieval pipelines.