Threat · curated 18 Aug 2026
Copilot tricked into telling reseachers how to hack itself
First reported theregister.com
Coverage timeline
Single-source incident — first reported, latest, and curated coincide.
Why it matters
CoSnitch shows that an AI assistant can be manipulated into disclosing its own security controls and hidden parameters during normal conversation, turning the model itself into an attack surface for data exfiltration and memory poisoning.
Varonis Threat Labs disclosed a Microsoft Copilot Personal vulnerability they named "CoSnitch," using a technique they call "meta-hacking" that social-engineers the AI's reasoning engine into revealing how to attack itself. By repeatedly asking Copilot why an attack wouldn't work, researchers extracted internal parameters and configuration details—including an undocumented autorun=1 parameter enabling auto-execution of ?q=-supplied prompts—and were able to exfiltrate sensitive data to an external server and poison Copilot's persistent memory. Reported to Microsoft in December 2025, with a patch and CVE planned.