Research · curated 31 Aug 2026
Can LLMs Reliably Self-Report Adversarial Prefills, and How?
First reported arxiv.org
Coverage timeline
Single-source research — first reported, latest, and curated coincide.
Why it matters
Adversarial prefill attacks reliably steer LLMs into producing harmful content that the models then falsely claim as their own intended output, showing that LLM self-reports cannot be trusted as a safety check for detecting compromised responses.
A KAIST research paper, "Can LLMs Reliably Self-Report Adversarial Prefills, and How?", evaluates whether ten open-weight instruction-tuned LLMs (3B-70B) can recognize that a prior response was elicited by an adversarial prefill attack. Across four safety benchmarks no model reliably recognizes its own compromised outputs, claiming intent on prefilled responses at an average rate of 25.3%, and the introspective signal depends heavily on refusal-direction reasoning and probe framing; training to improve introspection counterintuitively raises attack success under prefill.