Research · curated 28 Sep 2026

Attacks That Score, Attacks That Transfer: Guardrail Asymmetry, the Confused Deputy, and Hidden-Evaluator Reconstruction in a Multi-Step Tool-Agent Security Benchmark by Uchechukwu Ajuzieogu :: SSRN

Coverage timeline

28 Sep 2026ssrn.com

Single-source research — first reported, latest, and curated coincide.

Why it matters

The study exposes structural blind spots in agent guardrails — a 'nothing-blockable' confused-deputy attack family that evades provenance-based defenses — and offers concrete defense recommendations relevant to anyone hardening tool-using LLM agents.

A monograph by Uchechukwu Ajuzieogu forensically reconstructs the 2026 Kaggle AI Agent Security — Multi-Step Tool Attacks competition, where 4,252 teams built multi-step prompt-injection attacks against a tool-using LLM agent. It documents four attack predicates (data exfiltration, destructive write, unauthorized-to-allowed, confused deputy), reverse-engineers both public and hidden guardrails, and shows that attack transfer to the private evaluator was a property of attack family — confused-deputy attacks transferred near 1:1 while marker-carrying exfiltration attacks were reliably blocked by a persistent-provenance policy.