Tool
vllm-sr/jailbreak-detection-dataset · Datasets at Hugging Face
Publication date unknown · Discovered huggingface.co
Page published
Publication date unknown · First observed: 8 Oct 2026
Coverage timeline
Single-source research — one report is available.
Why it matters
The vllm-sr/jailbreak-detection-dataset gives defenders ready-made labeled data to build and benchmark classifiers that flag jailbreak and prompt-injection attempts against deployed LLMs.
The vllm-sr/jailbreak-detection-dataset on Hugging Face is a downloadable, labeled dataset of roughly 4,130 prompts split into train, validation, and test sets, with each entry tagged as either 'jailbreak' or 'benign' (e.g., DAN/APOPHIS-style role-play jailbreaks, harmful instruction requests, and ordinary queries). The dataset is auto-converted to Parquet and intended for training or evaluating LLM jailbreak-detection classifiers.