Tool

vllm-sr/jailbreak-detection-dataset · Datasets at Hugging Face

Page published

Publication date unknown · First observed: 8 Oct 2026

Coverage timeline

8 Oct 2026huggingface.coobserved

Single-source research — one report is available.

Why it matters

The vllm-sr/jailbreak-detection-dataset gives defenders ready-made labeled data to build and benchmark classifiers that flag jailbreak and prompt-injection attempts against deployed LLMs.

The vllm-sr/jailbreak-detection-dataset on Hugging Face is a downloadable, labeled dataset of roughly 4,130 prompts split into train, validation, and test sets, with each entry tagged as either 'jailbreak' or 'benign' (e.g., DAN/APOPHIS-style role-play jailbreaks, harmful instruction requests, and ordinary queries). The dataset is auto-converted to Parquet and intended for training or evaluating LLM jailbreak-detection classifiers.