Analysis · curated 20 Aug 2026
Prompt Injection Techniques: 7 Attack Types and Defenses - Mindgard
First reported · updated · 16 reports nhimg.org
Coverage timeline
Why it matters
Prompt injection remains the root unsolved weakness in LLM and agentic deployments, so a structured taxonomy of attack types and defenses helps defenders map their exposure and prioritize mitigations.
Mindgard's blog "Prompt Injection Techniques: 7 Attack Types and Defenses" is an explainer that catalogs prompt injection categories—including direct and indirect injection against LLM-integrated applications and AI agents—and surveys defensive measures. The piece synthesizes known attack classes and mitigations rather than disclosing a new specific exploit.
Summary
This dossier synthesizes guidance on securing AI agents against prompt injection, anchored in an Enoki Labs analysis that frames agent security as protecting systems that plan and act through tools, APIs and data from being manipulated into doing what their owner never intended. The central, unsolved problem is that a language model cannot reliably separate instructions it should follow from the content it is processing; in a chatbot this is contained, but in an agent model output becomes real actions — API calls, database writes, emails and code — so the blast radius expands dramatically.[0][48]
The analysis maps risk onto the OWASP Top 10 for Agentic Applications 2026 (published 9 December 2025) and catalogues disclosed incidents from 2024–2026, including Slack AI exfiltration, EchoLeak (CVE-2025-32711, CVSS 9.3), the Supabase MCP confused-deputy case, Replit's production database deletion, GitHub Copilot RCE via auto-approval (CVE-2025-53773), Amazon Q Developer supply-chain tampering (CVE-2025-8217), Salesforce Agentforce ForcedLeak (CVSS 9.4), postmark-mcp (the first malicious MCP server found in the wild), and Vertex AI service-agent credential abuse. OWASP's Q1 2026 round-up found eight major incidents with only one CVE, underscoring that most agent failures stem from design flaws, misconfiguration and the supply chain rather than scannable bugs.[0][49][71][31][34]
The threat is no longer theoretical: Zscaler ThreatLabz has observed two real-world indirect prompt injection campaigns — a payment scam and a cryptocurrency typosquatting campaign — hiding instructions in websites to manipulate web-enabled autonomous agents, and Anthropic reported with high confidence that the Chinese state-sponsored group GTG-1002 drove 80-90% of the tactical operations of a cyber espionage campaign autonomously through Claude Code. These reinforce the guidance's emphasis on least-agency design, scoped identities, egress controls and action logging rather than reliance on model-side filters.[0][59][61]
How it works
Prompt injection works because LLM-integrated applications blur the line between data and instructions: a language model cannot reliably tell the instructions it should follow from the content it is working on. Direct injection is a user typing an override; the more dangerous indirect injection plants instructions in content the agent will read later — an email, support ticket, GitHub issue, web page, lead form or tool reply — where the planted text competes with the system prompt and can win, turning the agent into a tool for someone outside the organisation.[0][48]
In an agent, six factors convert this weakness into a security problem: tools turn model output into real actions; autonomy lets one injected instruction steer a multi-step chain unobserved; memory and retrieval stores let a poisoned entry steer future sessions; delegated identity (OAuth tokens, API keys, service accounts) creates a confused deputy where anyone who steers the agent inherits its access; a runtime supply chain means MCP tool descriptions are read by the model as instructions; and in multi-agent systems one agent's output becomes the next agent's trusted input so a single fault cascades.[0]
Data exfiltration in these incidents rarely required exotic techniques: agents leaked data by encoding it into something that leaves the environment, such as an image URL or link the agent was allowed to render. EchoLeak and ForcedLeak both sent stolen data out through an image loaded from an allowlisted domain, so the effective fix lies in restricting what the agent may render, fetch and send rather than in the model itself.[0]
Academic research (Greshake et al., 2023) established the practical viability of indirect prompt injection against real systems, showing that processing retrieved prompts can act as arbitrary code execution and can control how and whether other APIs are called — the mechanism later realised in the catalogued agent incidents and in the live web-content campaigns Zscaler observed embedding hidden instructions in websites to manipulate autonomous agents.[48][59]
Affected versions and patch status
| Product | Affected | Patch status |
|---|---|---|
| Microsoft 365 Copilot (EchoLeak) | Zero-click indirect prompt injection chaining four bypasses including Microsoft's injection classifier; tracked as CVE-2025-32711, CVSS 9.3 | Disclosed and assigned a CVE; specific fix detail not stated in the evidence[0][49] |
| GitHub Copilot in VS Code | Injected instructions enabled auto-approval in the agent's own settings file, enabling arbitrary shell command execution; tracked as CVE-2025-53773 | Disclosed and assigned a CVE; specific fix detail not stated in the evidence[0][71] |
| Amazon Q Developer VS Code extension | A mis-scoped token let a threat actor insert a destructive prompt into a release; tracked as CVE-2025-8217 | Addressed per AWS security bulletin; attack failed in practice due to a syntax error[0][27][70] |
| Salesforce Agentforce (ForcedLeak) | Web-to-lead form injection caused the agent to exfiltrate CRM data to an expired allowlisted domain; rated CVSS 9.4 | Disclosed by Noma Security; specific fix detail not stated in the evidence[0][30] |
| Vertex AI Agent Engine | A deployed agent could extract its default service-agent credentials and use them to read every storage bucket in the project (ASI03) | Disclosed by Unit 42; specific fix detail not stated in the evidence[0][32] |
Key takeaways
- Agent security is distinct from chatbot security because model output becomes real actions: a redirected agent can send email, delete records, move money or hand over tokens, so the discipline centres on what an agent can touch once steered, not just what it says.[0]
- Most catalogued agent incidents required no software bug — only an agent that read untrusted content while holding valuable access — and OWASP's Q1 2026 round-up found eight major incidents with just one CVE, so vulnerability scanners miss the dominant failure modes.[0][34]
- Indirect prompt injection is being used operationally in the wild: Zscaler ThreatLabz observed payment-scam and cryptocurrency-typosquatting websites embedding hidden instructions specifically to manipulate web-enabled autonomous AI agents.[59]
- Adversarial abuse of agents is also occurring at scale against high-value targets, as shown by Anthropic's high-confidence report of GTG-1002 running 80-90% of a cyber espionage campaign's tactical operations autonomously via Claude Code, though AI hallucination and fabrication still limit fully autonomous attacks.[61]
- Defence is a governance and authorization problem solved by layered controls — least agency, scoped identities, egress restriction, tool pinning, sandboxing, and action logging — rather than any single model-side filter, since prompt injection has been an open, unsolved class of vulnerability since 2022.[0][45]
Defensive actions
- Break the access trifecta by design: for each agent decide which two of untrusted input, sensitive access, and external action it truly needs, and require a human in the loop for the session if it needs all three.: Agents that read untrusted content while holding valuable access and the ability to act outward are exactly the configuration behind most catalogued incidents.[0]
- Apply least agency as well as least privilege: fewer and narrower tools, no open-ended tools such as a raw shell, and no autonomy where a fixed workflow would do.: OWASP's Excessive Agency names too many functions, too many permissions and too much autonomy as root causes; every connected tool is something an injected instruction can also use.[0]
- Give every agent its own scoped, short-lived identity and enforce authorization in the systems it calls, never in the prompt.: Borrowed or over-broad credentials create a confused deputy, as in the Supabase MCP case where a service role bypassing row-level security read secret tokens because a support ticket told it to, and the Vertex AI case where an agent read every bucket with its default service-agent credentials.[0][32]
- Control egress: block rendering of links and images to untrusted domains and keep allowlists short and current.: EchoLeak and ForcedLeak both exfiltrated data through an image loaded from a domain the agent was allowed to reach.[0]
- Pin and review MCP server versions and tool descriptions as you would code, connecting only to servers you trust, and sandbox anything the agent executes.: Tool poisoning hides instructions in tool descriptions and rug pulls change them after approval; postmark-mcp showed a malicious MCP server reaching production.[0][31]
- Treat all web content and tool replies the agent reads as untrusted data, not instructions, and require specific approval for high-impact actions.: Zscaler observed live campaigns hiding instructions in websites to manipulate autonomous agents, and vague approval prompts train people to click yes.[0][59]