Research · curated 28 Jul 2026
GAVEL: Towards Rule-Based Safety Through Activation Monitoring
First reported arxiv.org
Coverage timeline
Single-source research — first reported, latest, and curated coincide.
Why it matters
GAVEL offers defenders an interpretable, updatable way to inspect LLM internals and flag harmful actions like phishing generation without retraining models, advancing detection beyond black-box input/output analysis.
Researchers from Ben-Gurion University's Offensive AI Research Lab introduce GAVEL, a model-agnostic, rule-based activation-monitoring framework that models LLM internal activations as composable 'cognitive elements' (CEs) to detect unsafe behavior in real time, akin to Snort/YARA rulesets. Presented at ICLR 2026 and slated for Black Hat USA 2026, GAVEL is open-sourced along with GAVEL Studio, an interactive rule-authoring tool, with code and datasets on GitHub.