First reported arxiv.org
Research · latest
First reported arxiv.org
Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs
Researchers from the University of South Dakota and Yangzhou University introduce a mechanistic interpretability framework that diagnoses LLM jailbreaks and adversarial prompts by constructing and aligning internal computation (attribution) graphs for clean versus attacked prompts. The method decomposes internal reasoning into invariant, suppressed, and emergent structures, identifies recurring vulnerability motifs, and performs causal interventions that improve robustness across multiple open-source LLMs and jailbreak benchmarks. Details →How the wire is made
Poll & cluster
Internet is crawled for AI security news and near-duplicate coverage is embedded and grouped into durable items.
Curate
AI Agent filters for agentic-AI relevance, classifies and tags each item, scores severity for threats, and writes the summary.
Every item here is one machine-curated intelligence object, not a headline.
Read the wire for free. There is a small charge to ask the index questions.
The wire, open
The complete curated feed, no key required.
- GET /feed.xml — RSS 2.0, every item
- GET /api/items — read-only
The vector desk
Query the index by meaning, not just keyword.
- GET /api/items?tags=&minSeverity=&itemType=
- GET /api/search?q= — keyword
- GET /api/semantic?q= — vector