Research · curated 15 Aug 2026

SpatialJB: How Text Distribution Art Becomes the “Jailbreak Key” for LLM Guardrails

Coverage timeline

15 Aug 2026arxiv.orgprimary

Single-source research — first reported, latest, and curated coincide.

Why it matters

SpatialJB shows that widely relied-upon output guardrails like the OpenAI Moderation API can be systematically penetrated using spatial text distribution, undermining a common safety layer defenders deploy around LLMs.

SpatialJB is a jailbreak technique from researchers at Zhejiang University and collaborators that exploits Transformers' weakness to spatially structured text perturbations, disrupting output generation so harmful content bypasses output guardrails. Experiments report near-100% attack success rates and over 75% success even against the OpenAI Moderation API, with baseline defenses also proposed; a demo video and code are provided.