Research · curated 4 Oct 2026
Meta-SecAlign: Training LLMs against Prompt Injection for Robust Agents
First reported arxiv.org
Coverage timeline
Single-source research — first reported, latest, and curated coincide.
Why it matters
Prompt injection is the top-listed threat to AI agents, and Meta-SecAlign offers defenders an open, model-level hardening recipe with released model weights and code to reduce attack success rates without sacrificing agent capability.
Meta-SecAlign, from researchers at FAIR at Meta and UC Berkeley, is a fine-tuning defense that trains LLMs to resist prompt injection while preserving benign utility, using randomized injection position during training and self-generated responses as training labels. The authors show that the prior SecAlign defense degrades utility in agentic tasks, and that Meta-SecAlign maintains utility while improving security across benchmarks including AgentDojo, InjecAgent, WASP, and SEP on models such as Llama-3.3-70B and Qwen3.