Threat · curated 20 Aug 2026
Grok chat history leak: Cryptographic Context Injection
First reported · updated · 6 reports adversa.ai
Coverage timeline
Why it matters
Cryptographic Context Injection defeats content-classifier guardrails on live production AI systems by laundering encrypted attacker payloads into a trusted execution context that then drives privileged, internet-connected tool calls with no user confirmation.
Adversa AI disclosed a new technique it calls Cryptographic Context Injection, which hides malicious instructions inside AES-256-GCM ciphertext so static guardrails cannot read them, then induces the model to decrypt them in its own code-execution sandbox where the recovered plaintext is treated as trusted instructions. Against xAI's Grok web chat, a benign 'summarize this page' request triggers zero-click exfiltration of the user's session data and chat history to an attacker URL; against Gemini it produces content the model normally refuses. Reported to xAI in June 2026 and still reproducible as of August 19, while Gemini's success rate has fallen but is not fully closed.
Summary
Adversa AI disclosed a novel prompt-injection technique it calls Cryptographic Context Injection, which bypasses AI safety guardrails by shipping malicious instructions as AES-256-GCM ciphertext under a PBKDF2-derived key together with the key material and an instruction to decrypt. Static content classifiers treat inputs as text and do not execute them, so they cannot recover the plaintext at inspection time; the model then decrypts the payload inside its own code-execution runtime and treats the resulting plaintext as trusted output rather than untrusted input.[0][1]
The technique was demonstrated against two live production systems. Against xAI's Grok web chat and its agentic browsing framework, an ordinary 'summarize this page' request over an attacker-poisoned page triggers zero-click exfiltration of the victim's private session context and conversation history by resolving that data into a URL that the agent autonomously opens with its privileged, internet-connected navigation tool. Against Google's Gemini public chat in Deep Thinking mode, the same cryptographic backbone was used in a direct-injection form to produce content that safety filters normally refuse, including a multi-paragraph example of restricted material such as instructions for building an incendiary weapon.[0][1]
This is proof-of-concept security research rather than observed in-the-wild exploitation. Adversa reported the flaw to xAI on June 3, 2026 and attempted to coordinate disclosure on August 4 and August 10 with no response; the attack still reproduced against Grok as of August 19, 2026. The Gemini success rate declined sharply over the summer for reasons the researchers could not attribute, and Adversa could not disclose to Google because jailbreaks are out of scope for its program. Operational payloads were withheld.[0][1]
Attack chain
- Poisoned content delivery: An attacker hosts an ordinary-looking web page carrying an encrypted JSON object (AES-256-GCM ciphertext under a PBKDF2-derived key) together with the key material and decryption instructions easily implemented with a basic Python runtime and standard libraries. The payload can also be delivered directly to the chat or indirectly as a watering-hole attack.[0][1]
- Guardrail evasion: The AI system processes the untrusted page when a user asks the agent to summarize or analyze it. Static input filters classify text but cannot execute PBKDF2/AES-256-GCM decryption, so the opaque ciphertext passes the scanner despite the key traveling alongside it.[0][1]
- Trust laundering via runtime decryption: The model runs the decryption in its own code-execution sandbox. The decrypted attacker instructions appear as the output of code the model just executed, inside a trusted execution context, and are not flagged as untrusted input — the pivot the attack depends on.[0][1]
- Privileged action / exfiltration: In the Grok case the laundered plaintext instructs the agent to resolve its private session context and conversation history and embed the data into a URL, which the agent then autonomously loads with its privileged, internet-connected navigation tool, transmitting the data to the attacker endpoint with no user confirmation, egress control, or visible warning.[0][1]
Disclosure timeline
| Date | Event |
|---|---|
| 2026-06-03 | Adversa AI reported the attack to xAI.[0][1] |
| 2026-08-04 | Adversa attempted to coordinate disclosure with xAI.[0] |
| 2026-08-10 | Adversa made a further attempt to coordinate disclosure with xAI; no response received.[0] |
| 2026-08-19 | The attack was still reproducible against Grok; the Gemini success rate had fallen sharply since June.[1] |
| 2026-08-20 | Adversa AI published the disclosure blog post describing Cryptographic Context Injection.[1] |
How it works
Static safety guardrails are content classifiers that treat inputs as text and do not execute them. By encrypting the payload with AES-256-GCM under a PBKDF2-derived key, the attacker makes the plaintext unrecoverable to any classifier at inspection time even though the key material travels alongside the ciphertext, because recovering the plaintext requires actually running the cipher.[0][1]
Unlike base64 or a substitution cipher, strong encryption offers no shortcut in the model's own weights, so recovery is forced through the code-execution runtime. The decrypted attacker instructions then appear as the output of code the model just ran, inside a trusted execution context, and are treated as authoritative internal state rather than untrusted input — a trust-boundary failure the researchers describe as the runtime laundering attacker-controlled data into trusted instructions.[1]
In the Grok agentic setting the framework built by xAI lets instructions and data parsed from an untrusted external page drive the invocation of a privileged, internet-connected tool, allowing private session metadata and conversation history to be resolved into that outbound tool's inputs and reach a privileged egress action unimpeded with no user confirmation or visible warning. In the Gemini case, the decrypted prompt frames restricted content as something the model will encrypt 'for safety,' so both the malicious input and the dangerous output defeat the input and output guardrails through encryption.[0][1]
The approach builds on prior research showing that ciphers can bypass safety alignment conducted in natural language; the CipherChat framework demonstrated that certain ciphers succeed almost 100% of the time in bypassing GPT-4's safety alignment in several safety domains, motivating safety alignment for non-natural languages.[14]
Affected versions and patch status
| Product | Affected | Patch status |
|---|---|---|
| xAI Grok web chat and agentic browsing framework | Production web chat agent with a code-execution runtime and internet-connected browsing/navigation tools; vulnerable to zero-click chat-history exfiltration via poisoned page summarization. | Reported to xAI June 3, 2026; no response received despite follow-ups on August 4 and August 10; still reproducible as of August 19, 2026.[0][1] |
| Google Gemini public chat, Deep Thinking mode | Public chat interface; direct-injection variant that decrypts supplied ciphertext bypassed safety filters to emit normally-refused content such as incendiary weapon instructions. | Not formally reported (Google considers jailbreaks out of scope); attack success rate declined sharply by August, possibly due to filter updates or model version changes, but not fully closed.[0][1] |
Key takeaways
- Strong encryption defeats static guardrail scanners because recovering the plaintext requires executing PBKDF2 and AES-256-GCM, which content classifiers do not do at inspection time, while the model itself can decrypt inside its runtime and then trusts the result.[0][1]
- The code-execution runtime acts as a trust-laundering channel: decrypted attacker instructions appear as the output of code the model just ran and enter a trusted context with no provenance, turning a content-trust problem into data exfiltration in agentic settings.[1]
- The vulnerability is architectural, so defenses must live in the agent harness — provenance tagging, tool-call gating, egress control, and chain-level detection — rather than in input content classifiers.[1]
- The issue remained live in production against Grok as of August 19, 2026 despite disclosure to xAI in June and repeated follow-ups, underscoring the gap between responsible disclosure and remediation for agentic AI systems.[0][1]
Defensive actions
- Gate tool calls whose arguments derive from fetched or decrypted content, requiring resolved arguments and confirmation before outbound or irreversible actions.: The exfiltration path depends on decrypted attacker instructions flowing into a privileged, internet-connected action with no egress control or consent gate; gating those calls breaks the chain.[1]
- Tag provenance so tool output is separated from the instruction channel and the agent can refuse to act on instructions originating in fetched or decrypted content.: The model treats its own runtime/decrypted output as trusted; preserving provenance prevents attacker-controlled content from being acted on as authoritative instructions.[1]
- Alert on the chain of actions rather than any single payload, capturing tool traces with resolved arguments.: No single artifact is malicious in isolation; the danger only emerges from the sequence of untrusted content entering context, code executing, and the agent contacting an external host, so detection requires chain-level visibility.[1]