Threat · curated 18 Aug 2026

Copilot tricked into telling reseachers how to hack itself

Coverage timeline

18 Aug 2026theregister.com

Single-source incident — first reported, latest, and curated coincide.

Why it matters

CoSnitch shows that an AI assistant can be manipulated into disclosing its own security controls and hidden parameters during normal conversation, turning the model itself into an attack surface for data exfiltration and memory poisoning.

Varonis Threat Labs disclosed a Microsoft Copilot Personal vulnerability they named "CoSnitch," using a technique they call "meta-hacking" that social-engineers the AI's reasoning engine into revealing how to attack itself. By repeatedly asking Copilot why an attack wouldn't work, researchers extracted internal parameters and configuration details—including an undocumented autorun=1 parameter enabling auto-execution of ?q=-supplied prompts—and were able to exfiltrate sensitive data to an external server and poison Copilot's persistent memory. Reported to Microsoft in December 2025, with a patch and CVE planned.