Tool · curated 26 Aug 2026

patronus-studio/wolf-defender-prompt-injection

Coverage timeline

25 Aug 2026huggingface.co

Single-source analysis — first reported, latest, and curated coincide.

Why it matters

Wolf Defender gives defenders a downloadable, deployable guardrail model to screen untrusted content for prompt injection and jailbreak attempts before it reaches tool-using LLM agents.

Wolf Defender is a multilingual ModernBERT-based (mmBERT-base) binary classifier published on Hugging Face by Patronus that detects prompt injections and jailbreak-style instructions before untrusted content reaches an LLM. The v2 release provides a 2,048-token context window, ONNX deployment variants, and benchmark results showing improved specificity on hard-benign inputs, and is intended as a local guardrail layer for AI agents, chatbots, and retrieval pipelines.