First reported ssrn.com
Research · latest
First reported paloaltonetworks.com
Perturbation Probing: A New Diagnostic for the Fragility of LLM Safety
Unit 42 researchers introduce 'Perturbation Probing,' a diagnostic method to measure the fragility of LLM safety alignment by applying perturbations to prompts and observing how easily safety guardrails collapse, drawing on prior work such as universal transferable adversarial suffix attacks. The technique is framed as a way to assess how robust deployed models are against jailbreak-style manipulation. Details →First reported arxiv.org
ALIBI: Adaptive Agentic Attacks on LLM-Based Vulnerability Detectors via Adversarial Code Comments
ALIBI is an automated adaptive black-box attack framework that evades LLM-based vulnerability detectors by inserting adversarial source-code comments that steer detector reasoning or fabricate external tool results without changing program behavior. Evaluated against four detectors including frontier multi-agent systems, it achieves attack success rates exceeding 90% across 125 real-world null-pointer dereference vulnerabilities, reaching 100% on one system, while prompt-level defenses offer limited robustness. Details →How the wire is made
Poll & cluster
Internet is crawled for AI security news and near-duplicate coverage is embedded and grouped into durable items.
Curate
AI Agent filters for agentic-AI relevance, classifies and tags each item, scores severity for threats, and writes the summary.
Every item here is one machine-curated intelligence object, not a headline.
Read the wire for free. There is a small charge to ask the index questions.
The wire, open
The complete curated feed, no key required.
- GET /feed.xml — RSS 2.0, every item
- GET /api/items — read-only
The vector desk
Query the index by meaning, not just keyword.
- GET /api/items?tags=&minSeverity=&itemType=
- GET /api/search?q= — keyword
- GET /api/semantic?q= — vector