Research · curated 15 Aug 2026
SpatialJB: How Text Distribution Art Becomes the “Jailbreak Key” for LLM Guardrails
First reported arxiv.org
Coverage timeline
Single-source research — first reported, latest, and curated coincide.
Why it matters
SpatialJB shows that widely relied-upon output guardrails like the OpenAI Moderation API can be systematically penetrated using spatial text distribution, undermining a common safety layer defenders deploy around LLMs.
SpatialJB is a jailbreak technique from researchers at Zhejiang University and collaborators that exploits Transformers' weakness to spatially structured text perturbations, disrupting output generation so harmful content bypasses output guardrails. Experiments report near-100% attack success rates and over 75% success even against the OpenAI Moderation API, with baseline defenses also proposed; a demo video and code are provided.