Frontier artificial intelligence models rely on intermediate reasoning steps that are deliberately withheld from end users to protect proprietary capabilities. When adversarial actors extract these hidden chains of thought, they can train smaller models on the full logic of leading systems rather than just their final answers. According to OpenAI, the company detected and shut down an adversarial distillation campaign targeting the protected reasoning of its AI models, which began at low volume on July 1 before escalating sharply.
As stated by OpenAI, the extraction effort spiked on July 24 and 25 to 16,000 requests originating from more than 4,000 users who relied on a shared query pattern. A deeper technical investigation uncovered a broader network of more than 15,000 accounts showing related extraction signatures, which OpenAI reported it fully shut down by July 28. OpenAI links a core group behind the activity to individuals associated with Moonshot AI, the company that builds the Kimi language model.
Shared Keys and Virtual Notepads Expose Logic
The mechanics of the breach exploited how reasoning data packets are encrypted and handled across customer sessions. AI providers routinely return hidden reasoning to clients as encrypted tokens that must be passed back with subsequent prompts to maintain context. As researcher Joachim Schaeffer and his team demonstrated in their paper, encrypted reasoning can be decrypted and written out using a separate conversation because those packets rely on shared encryption keys across model families.
This structural flaw allowed cheaper, smaller models from the same vendor family to act as decryption engines that printed the larger models' hidden thoughts word for word. It turns out that multi-billion-dollar R&D budgets can be bypassed if your encryption keys are essentially shared across the entire product lineup.
Cloud Endpoints Lagged Behind Primary Safeguards
Fixing vulnerabilities on a model provider's proprietary API does not automatically secure third-party distribution channels. When researcher Joachim Schaeffer and his team tested on September 13, the attack was blocked on OpenAI's and Anthropic's own interfaces, but it still worked on Microsoft Azure against every OpenAI model tried, including GPT-4.
According to the researchers' timeline, OpenAI did not add safeguards to the Azure endpoint until September 27. For Anthropic models, the reported extraction could no longer be reproduced on Azure starting September 28. As OpenAI explained, the company has banned fraudulent accounts, tightened sign-ups, and closed the hole that let people reuse and read encrypted reasoning. This multi-week lag between native API fixes and cloud partner patches highlights a glaring security blind spot in enterprise distribution.
Audit third-party cloud model deployments immediately to ensure provider-level safety patches and reasoning redactions are synchronized across all enterprise endpoints, unless you enjoy paying top-tier prices for models that are essentially leaking their own source code through the back door.