Threat · curated 1 Oct 2026

Disrupting a coordinated model-distillation campaign

Dossier

Coverage timeline

30 Sep 2026openai.comprimary

Single-source incident — first reported, latest, and curated coincide.

Why it matters

The attack shows that encrypted client-side reasoning traces are interchangeable across sessions and models within a provider, letting adversaries decode hidden chain-of-thought, exfiltrate PII and credentials, and bypass anti-distillation and safety safeguards at scale across OpenAI, Anthropic, and Google.

OpenAI reported disrupting a coordinated adversarial-distillation campaign that extracted protected chain-of-thought reasoning from its models by copying encrypted reasoning blocks from one conversation and asking a weaker same-provider model to decrypt and transcribe them. Activity began July 1, 2026, spiked to 16,000 requests from over 4,000 users on July 24–25, and was disrupted by July 28; OpenAI attributes a core cluster to individuals associated with Moonshot AI (developer of Kimi). Independent researchers (arXiv:2608.09867) separately documented the cross-model encrypted-reasoning interchangeability flaw enabling decryption jailbreaks, PII/credential recovery, and invisible prompt injection.

campaign

Summary

OpenAI detected and disrupted a coordinated adversarial-distillation campaign aimed at extracting protected model reasoning. The operators did not break encryption or access stored conversations; instead they manipulated model interactions so that hidden reasoning could be reproduced in forms visible to the requester, at scale and in violation of OpenAI's terms of service.[0]

The activity began July 1, 2026 at low volume, spiked on July 24-25 with roughly 16,000 attempted extraction requests from over 4,000 users, and was tied to a broader cluster of more than 15,000 users that OpenAI fully disrupted by July 28. OpenAI attributes a core cluster of this activity to individuals associated with Moonshot AI, the developer of Kimi.[0]

The campaign aligns with an architectural weakness documented by independent researchers: encrypted chain-of-thought blocks returned to clients are interchangeable across sessions, users, and models within a provider's ecosystem, letting an attacker inject a stronger model's encrypted trace into a weaker, less-safeguarded model and force it to output the reasoning verbatim in plaintext.[0][1]

Attack chain

  1. Capture encrypted reasoning: Operators obtained encrypted reasoning (chain-of-thought) blocks that models return to clients, including blocks copied from one conversation; researchers also showed such blocks could be scraped from publicly shared session logs.[0][1]
  2. Cross-model/cross-session replay: Because encrypted reasoning blocks are compatible and interchangeable across sessions, users, and models within a provider ecosystem, the captured trace was injected into another conversation or a weaker, less-safeguarded model from the same provider.[0][1]
  3. Forced decryption and transcription: The receiving model was prompted to decrypt and transcribe the hidden reasoning verbatim in plaintext, exposing content withheld from the final answer without directly jailbreaking the more capable model.[0][1]
  4. Scaled, coordinated extraction: The extraction pattern was executed in a coordinated, high-volume manner — about 16,000 requests from over 4,000 users during the July 24-25 spike, tied to a cluster exceeding 15,000 users — consistent with systematic adversarial distillation.[0]

Disclosure timeline

DateEvent
2026-07-01Coordinated extraction activity began at low volume.[0]
2026-07-24 to 2026-07-25High-volume spikes of about 16,000 attempted extraction requests from over 4,000 users using a relevant extraction pattern.[0]
2026-07-28OpenAI fully disrupted the related prompt-pattern activity across a cluster of more than 15,000 users.[0]
2026-08-10Researchers submitted the arXiv paper 'Stealing Reasoning Traces from Proprietary LLM APIs' documenting the underlying vulnerability.[1]
2026-09-30OpenAI published its account of disrupting the coordinated model-distillation campaign.[0]

Actor profile

Individuals associated with Moonshot AI

OpenAI states it is unclear whether all observed operators originated from a single actor, but it attributes a core cluster of the activity to individuals associated with Moonshot AI, the developer of the Kimi model. The attribution is OpenAI's assessment and is single-sourced.[0]

How it works

Providers conceal step-by-step reasoning and, rather than storing traces server-side, return them to the client as blocks of encrypted text that the client passes back with each request. Researchers found these encrypted blocks are fully compatible and interchangeable across different sessions, users, and models within a provider's ecosystem.[1]

By injecting an encrypted reasoning trace from a capable model into a weaker, less-safeguarded model from the same provider, an attacker forces the weaker model to decode and output the trace verbatim in plaintext — circumventing anti-distillation controls without directly jailbreaking the stronger model. OpenAI confirmed this replay path was real and closed the pathway that let someone who already possessed another user's encrypted reasoning replay it and recover its contents.[0][1]

The flaw enables four distinct attack vectors: circumventing anti-distillation mechanisms, large-scale private data extraction, revelation of hazardous information hidden in reasoning even when the final output safely refuses, and invisible prompt injection via payloads embedded in encrypted blocks to poison public agentic rollouts.[1]

Affected versions and patch status

ProductAffectedPatch status
OpenAI models (protected/hidden reasoning)Models returning client-side encrypted reasoning blocks; replay/decryption pathway exploited in the wildOpenAI deployed mitigations: closed the replay pathway, added checks to detect and hold streamed output exposing reasoning, strengthened protections across users, workspaces, organizations, and model families; further work ongoing.[0]
Anthropic and Google frontier modelsResearchers demonstrated the cross-model encrypted-reasoning extraction technique against Anthropic, OpenAI, and Google systems; described as a shared, non-provider-unique security challenge.Researchers proposed cryptographic and system-level mitigations following responsible disclosure; OpenAI shared findings via the Frontier Model Forum.[0][1]

Key takeaways

  • Adversarial distillation via client-side encrypted reasoning is a cross-industry architectural problem, not an OpenAI-specific vulnerability, affecting multiple frontier providers.[0][1]
  • Returning encrypted chain-of-thought to clients that is interchangeable across sessions, users, and models creates a replay-and-decrypt path that bypasses anti-distillation and safety controls without directly jailbreaking the strong model.[1]
  • Defending against scaled distillation requires layered, adaptive controls — technical protections, coordinated detection/enforcement against campaigns, and cross-industry and government information sharing.[0]

Defensive actions

  • Prevent replay of client-held encrypted reasoning across sessions, users, and models: The attack relies on encrypted reasoning blocks being interchangeable within an ecosystem; binding reasoning artifacts to their originating session/user/model closes the replay-and-decrypt pathway OpenAI remediated.[0][1]
  • Inspect and hold streamed tool/model output that may expose hidden reasoning: OpenAI added checks to detect and hold streamed output that might expose reasoning; tool-output attacks require protections that examine more than ordinary visible text.[0]
  • Monitor for coordinated high-volume extraction prompt patterns and enforce account controls: OpenAI banned/restricted fraudulent accounts, strengthened signup and infrastructure controls, and expanded monitoring for related networks after detecting a prompt-pattern cluster of 15,000+ users.[0]
  • Treat publicly shared session logs as sensitive and avoid exposing encrypted reasoning blocks: Researchers recovered 367 PII artifacts and 182 credentials by decoding 315,320 reasoning blocks scraped from public repositories, showing shared logs can leak hidden contents.[1]
  • Extend reasoning protections to partner-hosted and third-party deployments: OpenAI notes partner-hosted deployments need the same protections as first-party services, and it coordinated with third-party providers to disrupt involved accounts.[0]