Research · curated 11 Sep 2026

Cross-Layer Manifold Attack and Projection for Robust Safety Alignment of Large Language Models by Ziwen Peng, Jiayu Du, Qi Zhou, Jin Zhu, Jianpeng Li :: SSRN

Coverage timeline

11 Sep 2026ssrn.com

Single-source research — first reported, latest, and curated coincide.

Why it matters

CMA demonstrates that aligned LLMs retain latent unsafe pathways reactivatable by adversarial prompts, showing defenders that output-level refusal is insufficient and that representation-level hardening is needed against jailbreaks.

A preprint by Ziwen Peng and colleagues at Purple Mountain Laboratories proposes Cross-layer Manifold Attack (CMA), which manipulates hidden-state representations across layers to shift an LLM from refusal toward compliance and induce jailbreak responses, and Cross-layer Manifold Projection (CMP), a defensive safety-tuning method that hardens models against such latent-space perturbations. Experiments across multiple model architectures show CMP improves robustness against adversarial jailbreaks while preserving general capabilities and reducing over-refusal.