First reported trendmicro.com
Lead dispatch
First reported · updated · 3 reports embracethered.com
AWS Kiro: Arbitrary Code Execution via Indirect Prompt Injection
Researchers found a vulnerability (CVE-2026-10591) in AWS Kiro, an agentic IDE, where hidden instructions planted in a web page or source file that Kiro processes can trigger indirect prompt injection to rewrite Kiro's own MCP server configuration (~/.kiro/settings/mcp.json) or allowlist arbitrary Bash commands in .vscode/settings.json, achieving arbitrary code execution on the developer's machine with no approval prompt. The human-in-the-loop approval boundary is bypassed because Kiro can write to these config files without user consent, and AWS has issued a fix and CVE.indirect-prompt-injection · prompt-injection · remote-code-execution · tool-abuse · config-poisoning
ai-agents · mcp · llm · agentic-ide
The wire · latest
First reported simonwillison.net
Understanding ChatGPT Work
Simon Willison's teardown of OpenAI's ChatGPT Work (specifically the cloud variant, Work Cloud) argues its feature set — internet-enabled code execution, a headless Chrome browser, a persistent scratch filesystem, sub-agents, scheduled automations, and Cloudflare Workers site deploys — combines all three elements of his 'lethal trifecta': access to private data, exposure to untrusted content, and a channel to exfiltrate stolen data. Willison does not demonstrate an exploit but asks OpenAI to explain how it defends Work sessions against prompt injection, criticizing the product's opacity around system prompts and tool descriptions. Details →First reported mdazlaanzubair.com
Are Leaked System Prompts Helpful?
An analysis piece by Muhammad Azlaan Zubair examines a GitHub repository (asgeirtj/system_prompts_leaks) with 63K+ stars collecting supposedly leaked system prompts from ChatGPT, Claude, Gemini, Grok and others, questioning whether the exposures are truly accidental and noting some may have been extracted via prompt injection. Rather than resolving provenance, the author argues the prompts' layered structure reveals production LLM application architecture patterns like tool contracts, routing rules, and permission behavior. Details →First reported mindgard.ai
Bypassing ChatGPT Image Safeguards Through Memory Manipulation
Mindgard research demonstrates bypassing ChatGPT's image-generation safeguards through manipulation of custom memory and system/instruction context, inducing policy-inconsistent output including sexualized images of fictitious and real people. The techniques exploit the bio tool, model set context, and image routing/filtering pipeline without accessing model weights, and were disclosed to OpenAI prior to publication. Details →First reported · updated · 4 reports openai.com
Disrupting a new covert influence campaign from Russia
OpenAI banned a cluster of ChatGPT accounts originating in Russia that were used to generate English-language social media comments across Substack, Telegram, X, Facebook and LinkedIn to promote the International Burke Institute (IBI), a front presenting itself as an Israel-based 'expert community.' The operators used VPNs to bypass Russia access restrictions and instructed ChatGPT to hide linguistic clues of their Russian origin, as part of a covert influence campaign that reached relatively small audiences. Details →First reported openai.com
Disrupting a new covert influence campaign from Russia
OpenAI reported banning a cluster of ChatGPT accounts very likely originating in Russia that used the model to generate English-language social media comments promoting the International Burke Institute, a covert influence operation that hid its Russian origins using VPNs and prompted ChatGPT to mask linguistic tells. The operation combined AI-generated posts across Substack, Telegram, X, Facebook, and LinkedIn with a website of copied and misattributed academic work and a pro-Russia 'sovereignty' index. Details →First reported theregister.com
OpenAI ditches Recall-style screenshot surveillance for friendly keylogging
OpenAI's new opt-in 'Computer History' feature for the ChatGPT macOS desktop app captures user interaction events (clicks, typing, keyboard shortcuts, app switches) via macOS accessibility APIs, turning them into text summaries and local memory files to build ChatGPT memories. The Register notes the files are stored unencrypted locally for up to 48 hours, are accessible to other programs running as the same user, and increase the user's exposure to prompt injection. Details →First reported · updated · 2 reports hix.ai
ChatGPT No Restrictions (Ultimate Guide for 2026) | God of Prompt
A how-to guide titled 'How to Jailbreak ChatGPT' walks readers through several well-known jailbreak techniques against ChatGPT, including the 'DAN' (Do Anything Now) persona, a 'Developer Mode' simulation, and a hypothetical narrative frame, and supplies sample prompts intended to bypass OpenAI's safety alignment. The piece also lists risks such as account suspension, exposure to harmful content, and increased hallucinations. Details →First reported darkreading.com
Researcher Claims Control of ChatGPT Secure Sandbox
At Black Hat USA 2026, Palo Alto Networks researcher Simcha Kosman presented "A Billion-User Blast Radius: Owning ChatGPT's Secure Sandbox," a proof-of-concept attack chain that bypasses ChatGPT's LLM supervisor to achieve persistent root execution inside its isolated container sandbox, establishing C2-style control. The demonstration showed how a victim's ChatGPT session could be tricked into escaping the runtime's intended controls, though it is a PoC rather than an attack against a realistic enterprise environment. Details →First reported bbc.com
OpenAI works to stop ChatGPT generating 'sex crime scene' images
Researchers at Mindgard demonstrated that a simple, slightly-altered prompt jailbreaks SpaceXAI's Grok (and previously OpenAI's ChatGPT/GPT-5.4) into generating graphic sexual and violent images without explicitly requesting such content. The same technique could be adapted to produce deepfakes of real people; OpenAI added safeguards after disclosure but researchers say small changes still bypass them. Details →First reported medium.com
AI Jailbreaking Is Now Sold as a Service. Here’s What That Means for You | by Muhammad Haider Tallal | MeetCyber
An analysis piece describes the rise of "jailbreak-as-a-service," where dark web vendors sell subscription frameworks (reportedly around $75/month) that coerce commercial LLMs like Claude into writing malware, lowering the barrier for low-skill actors. The article contrasts this with earlier bespoke criminal models such as WormGPT and FraudGPT. Details →First reported paloaltonetworks.com
Analyzing the Current State of AI Use in Malware
Palo Alto Networks Unit 42 published a threat-research analysis examining how malware authors are currently incorporating generative AI and LLMs (such as ChatGPT) into their tooling, referencing observed samples including infostealers and Sliver-based implants. The piece assesses the practical maturity and limitations of AI use in real-world malware based on analyzed artifacts. Details →First reported · updated · 3 reports zenity.io
ChatGPT AgentForger Flaw Could Deploy Rogue Workspace Agents via a Phishing Link
Zenity Labs disclosed "AgentForger," a flaw in OpenAI's ChatGPT workspace agent builder that let a single crafted ChatGPT link silently create, configure, publish, and schedule an attacker-controlled autonomous agent inside a victim's workspace. The proof-of-concept agent inherited the employee's identity and connected apps (Outlook, Teams, Slack, SharePoint, Google Drive), disabled approval prompts, and used inbox messages tagged "TASK" as a covert command-and-control channel to search and exfiltrate corporate data. Details →How the wire is made
Poll & cluster
Internet is crawled for AI security news and near-duplicate coverage is embedded and grouped into durable items.
Curate
AI Agent filters for agentic-AI relevance, classifies and tags each item, scores severity for threats, and writes the summary.
Every item here is one machine-curated intelligence object, not a headline.
Read the wire for free. There is a small charge to ask the index questions.
The wire, open
The complete curated feed, no key required.
- GET /feed.xml — RSS 2.0, every item
- GET /api/items — read-only
The vector desk
Query the index by meaning, not just keyword.
- GET /api/items?tags=&minSeverity=&itemType=
- GET /api/search?q= — keyword
- GET /api/semantic?q= — vector