First reported ssrn.com
Lead dispatch
First reported · updated · 3 reports embracethered.com
AWS Kiro: Arbitrary Code Execution via Indirect Prompt Injection
Researchers found a vulnerability (CVE-2026-10591) in AWS Kiro, an agentic IDE, where hidden instructions planted in a web page or source file that Kiro processes can trigger indirect prompt injection to rewrite Kiro's own MCP server configuration (~/.kiro/settings/mcp.json) or allowlist arbitrary Bash commands in .vscode/settings.json, achieving arbitrary code execution on the developer's machine with no approval prompt. The human-in-the-loop approval boundary is bypassed because Kiro can write to these config files without user consent, and AWS has issued a fix and CVE.indirect-prompt-injection · prompt-injection · remote-code-execution · tool-abuse · config-poisoning
ai-agents · mcp · llm · agentic-ide
The wire · latest
First reported paloaltonetworks.com
Perturbation Probing: A New Diagnostic for the Fragility of LLM Safety
Unit 42 researchers introduce 'Perturbation Probing,' a diagnostic method to measure the fragility of LLM safety alignment by applying perturbations to prompts and observing how easily safety guardrails collapse, drawing on prior work such as universal transferable adversarial suffix attacks. The technique is framed as a way to assess how robust deployed models are against jailbreak-style manipulation. Details →First reported sparai.org
Evading Detection in LLM Jailbreaking - SPAR Project
A SPAR research project proposal led by Leo Schwinn (TU Munich/Helmholtz) outlines a novel jailbreak method that optimizes adversarial attacks (suffix or refusal-direction objectives) strictly on benign over-refusals, then tests whether they transfer to harmful tasks — the goal being attacks built without ever touching harmful content, thereby evading content classifiers and provider monitoring. A linked companion paper argues LLM-as-a-Judge safety evaluators degrade to near-random reliability under adversarial distribution shifts, inflating reported attack success rates. Details →First reported arxiv.org
ALIBI: Adaptive Agentic Attacks on LLM-Based Vulnerability Detectors via Adversarial Code Comments
ALIBI is an automated adaptive black-box attack framework that evades LLM-based vulnerability detectors by inserting adversarial source-code comments that steer detector reasoning or fabricate external tool results without changing program behavior. Evaluated against four detectors including frontier multi-agent systems, it achieves attack success rates exceeding 90% across 125 real-world null-pointer dereference vulnerabilities, reaching 100% on one system, while prompt-level defenses offer limited robustness. Details →How the wire is made
Poll & cluster
Internet is crawled for AI security news and near-duplicate coverage is embedded and grouped into durable items.
Curate
AI Agent filters for agentic-AI relevance, classifies and tags each item, scores severity for threats, and writes the summary.
Every item here is one machine-curated intelligence object, not a headline.
Read the wire for free. There is a small charge to ask the index questions.
The wire, open
The complete curated feed, no key required.
- GET /feed.xml — RSS 2.0, every item
- GET /api/items — read-only
The vector desk
Query the index by meaning, not just keyword.
- GET /api/items?tags=&minSeverity=&itemType=
- GET /api/search?q= — keyword
- GET /api/semantic?q= — vector