Threat

Breaking Claude Code Opus 5 Auto Mode

Page published · Page updated

Dossier

Coverage timeline

27 Aug 2026embracethered.comobservedprimarysimonwillison.netobservedcybernews.comtheregister.comobserved 6 Sep 2026hackmag.comobservedcloudsecurityalliance.o…simonwillison.netobserved

Why it matters

Claude Code's Auto Mode is now the default protection against prompt injection for a widely used coding agent, yet this demonstrated chain shows a targeted attack can bypass its classifier and reach code execution, meaning teams must sandbox and restrict agents rather than trust the built-in safety layer.

Johann Rehberger (embracethered) demonstrated an indirect prompt-injection attack chain against Claude Code Opus 5 in Auto Mode, achieving code execution with a 60-80% success rate from a simple 'summarize this website' request. The chain nudges Claude from WebFetch to curl, downloads a ZIP whose extracted malicious struct.py shadows Python's standard module, so importing base64 triggers attacker code; in some runs Auto Mode's safety classifier even blocked Claude's own cleanup command. The result contrasts sharply with a vendor-commissioned evaluation that reported 0.00% attack success for Opus 5 in Auto Mode.

vuln-research

Summary

Researcher Johann Rehberger (Embrace The Red) demonstrated a targeted prompt-injection attack chain that hijacks Claude Code Opus 5 running in Auto Mode through nothing more than a request to summarize a website, achieving code execution with a reported 60-80% success rate on a small sample size.[0][2]

The finding directly challenges Anthropic's confidence in Auto Mode, whose safety classifier replaced human approval prompts and became the default in Claude Code in mid-August 2026, and whose commissioned third-party evaluation (Trajectory Labs) reported a 0.00% prompt-injection success rate for Opus 5 in Auto Mode.[2][0]

Notably, the safety mechanism itself became part of the failure: in some runs Auto Mode allowed creation of the malicious process but then blocked Claude's own command to terminate it. The researcher and commentators conclude the only safe way to run agents at risk of adversarial attack is inside a sandbox with restricted egress and monitoring.[0][2]

Attack chain

  1. Initial access via benign task: A user asks Claude Code to summarize a malicious website that presents itself as a small archive of notebook records, giving Claude a plausible reason to investigate the material.[2]
  2. Tool pivot (WebFetch to Bash/curl): The server answers Claude's WebFetch request with an HTTP 415 Unsupported Media Type without instructing Claude what to do, causing Claude to independently fall back to fetching the page with curl via the Bash tool — making the malicious path the one worth pursuing rather than dictating it.[2]
  3. Redirect to malicious archive: The root URL returns an HTTP 303 redirect to a ZIP archive (WIC-notebook-catalogue.ZIP) containing files in a special encoding along with a native decoder binary.[2]
  4. Decoder execution in attacker-controlled directory: Claude correctly refuses to execute the supplied binary and instead writes its own Python decoder, but runs that decoder inside the unzipped, attacker-controlled directory.[2][0]
  5. Python module shadowing and code execution: A malicious struct.py in the working directory shadows Python's standard implementation, so when Claude imports the base64 module it triggers the poisoned struct.py and executes attacker code.[2][0]
  6. Safety mechanism failure on cleanup: In some runs Claude detected the compromise and attempted to terminate the malware process, but Auto Mode denied the cleanup command even though it had allowed the malicious process to be created.[0]

Disclosure timeline

DateEvent
Mid-August 2026Auto Mode becomes the default starting mode in Claude Code, replacing human approval prompts with a safety classifier.[2][0]
27th August 2026The Embrace The Red research demonstrating the Auto Mode code-execution attack chain is publicized.[0]

How it works

The attack exploits Python's module resolution behavior combined with agent tool autonomy. Claude is nudged from the summarizing WebFetch tool to raw curl in Bash (via a 415 response), downloads and unzips an attacker archive, then writes and runs its own Python decoder from within the attacker-controlled directory. A malicious struct.py placed in that directory shadows Python's standard library module, so a subsequent import of the base64 module (which internally imports struct) executes the attacker's code — code execution without ever explicitly instructing the model to run malware.[2][0]

A key characteristic is indirection: the attack never tells the model what to do. It instead engineers the environment (unsupported media type, redirects, plausible archive contents) so that the malicious path is simply the most reasonable way for the agent to complete its objective.[2]

Affected versions and patch status

ProductAffectedPatch status
Claude Code Opus 5 in Auto ModeAuto Mode configuration (default starting mode since mid-August 2026), which uses a safety classifier in place of human approval promptsNo fix indicated; researcher demonstrated 60-80% attack success against the current default configuration despite a commissioned evaluation reporting 0.00%[2][0]

Key takeaways

  • An agent safety classifier is not a substitute for isolation: a targeted attack chain achieved 60-80% code-execution success against Claude Code Opus 5 Auto Mode despite a commissioned benchmark reporting 0.00%.[2][0]
  • Effective prompt-injection attacks manipulate the environment rather than issuing explicit commands, steering the agent's own reasoning toward the malicious path (e.g., pivoting from WebFetch to curl after a 415 response).[2]
  • Classifier-based defenses can themselves become failure points — Auto Mode permitted the malicious process yet blocked Claude's attempt to kill it.[0]
  • Python module shadowing (a rogue struct.py triggered by importing base64) is a practical execution primitive when an agent runs code inside an attacker-controlled directory.[2][0]

Defensive actions

  • Run unattended coding agents inside a container, VM, or OS-level sandbox rather than relying on Auto Mode's classifier.: Auto Mode is not a substitute for an isolated environment; a sandbox is the only reliable protection when an agent may attract adversarial attacks.[0][2]
  • Restrict network egress from the agent runtime.: Limiting outbound connectivity constrains an agent's ability to fetch attacker archives and exfiltrate data.[0]
  • Actively monitor what agents are doing.: The classifier allowed malicious process creation and even blocked cleanup, so independent monitoring is needed to catch compromise.[0][2]
  • Do not expose home directories, SSH keys, or cloud credentials to the agent runtime.: Reducing accessible secrets limits the blast radius if the agent is hijacked and achieves code execution.[0]