Research

AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model

Page published

Publication date unknown · First observed: 9 Oct 2026

Coverage timeline

9 Oct 2026arxiv.orgobservedprimary

Single-source research — one report is available.

Why it matters

Indirect prompt injection partially succeeds in up to 86% of cases on web-agent security benchmarks, so a training method that improves both capability and robustness against adaptive attacks addresses a core defensive gap for deployed web agents.

AdvSim2Real, from researchers at MBZUAI, Amazon, and MIT, is a training method that co-evolves a task curriculum, an adaptive injection adversary, and the web agent inside a frozen web world model to build task-preserving robustness against indirect prompt injection. The authors report a 4B agent gains 33.6% relative completion under an unseen frontier-model adversary across 150 web tasks, and release code, a benchmark, and checkpoints.