Research · curated 29 Jul 2026
The Joint Effect of Quantization and Sampling Temperature on LLM Safety Alignment: A Factorial Analysis
First reported arxiv.org
Coverage timeline
Single-source research — first reported, latest, and curated coincide.
Why it matters
The study shows that common deployment choices like sampling temperature can materially erode LLM safety alignment and that single-benchmark safety claims understate real jailbreak risk, informing how defenders evaluate deployed models.
A factorial study by Hari Prasad and Ritam Pal evaluates how weight quantization and sampling temperature jointly affect LLM safety alignment across 8 instruction-tuned models, 3 precisions, and 6 temperatures (144 configurations, ~2 million responses scored by a six-judge ensemble). The authors find standard INT4/INT8 quantization is roughly safety-neutral for 7 of 8 models, while higher sampling temperatures sharply increase decision instability (DFR up to 41.9% at T=1.0), and the two factors do not compound; single-benchmark evaluation substantially understates jailbreak risk.