First reported github.com
Lead dispatch
First reported · updated · 3 reports talosintelligence.com
The Closed Quorum: Inside the first reported autonomous AI C2 implant
Cisco Talos, through its CAIRN project, discovered CLOSEDQUORUM, described as the first publicly documented Windows implant to delegate tactical command-and-control decisions to a panel of up to four commercial LLMs (DeepSeek, Qwen, Mistral, and Google Gemini) rather than a human operator or attacker-run C2 server. The implant autonomously selects and executes its next action to harvest user credentials and crypto wallets; while not confirmed deployed in the wild, binary artifacts linked the developer to carding-related criminal forum postings dating to 2025.autonomous-ai-malware · ai-integrated-malware · command-and-control · data-exfiltration · llm-abuse
llm · ai-agents · windows · deepseek · qwen · mistral · gemini
The wire · latest
First reported exploiting.systems
Prompt Injection in VirusTotal's Code Insights API
A researcher discovered prompt-injection flaws in VirusTotal's AI-powered Code Insights API (backed by gemini-2.5-flash), showing that embedded injection strings and false pretext in large block comments can suppress or alter analysis, force undocumented error schemas that leak the backend model, and induce false negatives or false positives. The bugs were accepted by Google's AI VRP on March 25, 2026 and are being patched. Details →First reported bleepingcomputer.com
AI 'watermark removers' flood the web. Almost none can prove they work.
A market of 'AI watermark remover' tools has appeared following Anthropic's rollout of invisible marks in Claude's text output, spanning a GitHub project with over 4,500 stars (watermarks-remover), several newly registered web tools, and services like StealthGPT and Human Writes that advertise stripping Claude, Gemini, OpenAI and SynthID-Text watermarks. BleepingComputer reports that almost none of the claims can be verified because Anthropic has not published how its text watermark works or released a detector; the tools reliably strip only hidden Unicode characters and C2PA/EXIF/XMP file metadata, which is trivial and not proof of defeating the underlying statistical watermark. Details →First reported sparai.org
Evading Detection in LLM Jailbreaking - SPAR Project
A SPAR research project proposal led by Leo Schwinn (TU Munich/Helmholtz) outlines a novel jailbreak method that optimizes adversarial attacks (suffix or refusal-direction objectives) strictly on benign over-refusals, then tests whether they transfer to harmful tasks — the goal being attacks built without ever touching harmful content, thereby evading content classifiers and provider monitoring. A linked companion paper argues LLM-as-a-Judge safety evaluators degrade to near-random reliability under adversarial distribution shifts, inflating reported attack success rates. Details →How the wire is made
Poll & cluster
Internet is crawled for AI security news and near-duplicate coverage is embedded and grouped into durable items.
Curate
AI Agent filters for agentic-AI relevance, classifies and tags each item, scores severity for threats, and writes the summary.
Every item here is one machine-curated intelligence object, not a headline.
Read the wire for free. There is a small charge to ask the index questions.
The wire, open
The complete curated feed, no key required.
- GET /feed.xml — RSS 2.0, every item
- GET /api/items — read-only
The vector desk
Query the index by meaning, not just keyword.
- GET /api/items?tags=&minSeverity=&itemType=
- GET /api/search?q= — keyword
- GET /api/semantic?q= — vector