Research · curated 22 Sep 2026

Structural Jailbreaks Generalize but Do Not Compound: A cross-provider and multilingual study of Involuntary In-Context Learning

Coverage timeline

22 Sep 2026arxiv.orgprimary

Single-source research — first reported, latest, and curated coincide.

Why it matters

The IICL structural jailbreak defeats content-level safety by operating at the pattern-completion layer and generalizes across providers, hitting near-100% success on financial-abuse behaviors, showing defenders that structural framing attacks are a dominant residual risk for deployed aligned models.

A red-team study tests whether two known LLM weaknesses—Involuntary In-Context Learning (IICL), a structural jailbreak that reframes a harmful request as the final missing cell of a JSON data-labeling task, and the multilingual safety gap—compound when combined against Google Gemini models (gemini-2.5-flash and flash-lite). IICL generalized to Gemini and raised attack success to 80–90% on HarmBench and 97–100% on the FinProof financial-abuse benchmark, but forcing non-English output attenuated rather than stacked the attack (11 of 12 non-English conditions scored below their English baseline), attributed to a 'relevance curse' where models produce lower-quality harmful content in lower-resource languages.