OpenAI Says It Disrupted a Reasoning-Extraction Campaign Linked to Moonshot AI
The company attributes a core cluster to individuals associated with Kimi’s developer, not every operator. Its request counts measure attempts, not confirmed successful extractions.
OpenAI says it disrupted a campaign by July 28 that sought to expose protected model reasoning for use in training other systems, following 16,000 extraction-pattern requests from more than 4,000 users on July 24–25. A broader cluster involved related prompt activity from more than 15,000 users; OpenAI links only a core group to people associated with Moonshot AI, and says the figures do not establish successful extraction or a single operator. The incident led to changes in account controls and reasoning delivery, while defenses for partner-hosted deployments and tool-output attacks remain in progress.
01
The reported technique used one conversation to copy encrypted reasoning and another to coax the model into decrypting and transcribing it; OpenAI says attackers did not break encryption or access stored chats.
02
OpenAI closed a replay pathway that could let someone resubmit another user's encrypted reasoning and recover its contents.
03
The company also tightened signup, infrastructure and monitoring controls, added checks on streamed output, and worked with third-party providers to disrupt related accounts.
OpenAIdisclosed on September 30 that it had disrupted a coordinated effort to expose its models’ hidden reasoning for use in training other AI systems. It attributed a core cluster to individuals associated with Moonshot AI, the developer of Kimi. The alleged route was manipulated conversations—not a database breach—and OpenAI says the risk extends beyond its own models.
A quiet start, then a July surge
The activity began at low volume on July 1, 2026. OpenAI describes it as adversarial distillation: systematically using a model’s outputs or reasoning without authorization to train, reproduce or improve another model. The target was protected reasoning, the internal record a model uses to work through a task. That record can contain information withheld from its final answer.
One technique involved copying encrypted reasoning from one conversation, then asking a model in another conversation to decrypt and transcribe it. OpenAI says the operators did not break encryption or directly access stored user conversations. Instead, they manipulated interactions to make protected material visible to the requester, at a coordinated scale that violated its terms of service.
The surge arrived on July 24 and 25. OpenAI counted 16,000 requests using a relevant extraction pattern from more than 4,000 users. Its investigation then identified related prompt patterns across a larger cluster, which it says it fully disrupted by July 28. Those figures describe attempted extraction; they do not establish how often the techniques worked.
The wider cluster
More than 15,000Users with related prompt-pattern activity
OpenAI says it disrupted this cluster by July 28. The user count is not a count of successful extractions.
Closing the replay route
Independent security researchers also responsibly disclosed related weaknesses involving different models and conversation compaction, or condensed conversation histories. OpenAI says it confirmed those attack paths were real. Their findings helped it understand the wider class of attacks and speed up mitigations; the company does not attribute those researchers’ work to the campaign.
OpenAI’s response combined account enforcement with changes to how reasoning is protected and delivered. It says it strengthened protections across users, workspaces, organizations and model families. A specific fix closed a replay pathway: someone who already possessed another user’s encrypted reasoning could submit it again and recover its contents.
Other defenses OpenAI deployed
Banned or restricted fraudulent accounts, tightened signup and infrastructure controls, and expanded monitoring for related networks.
Added checks to detect and hold streamed output that might expose protected reasoning.
Worked with third-party service providers to identify and disrupt accounts when related activity moved through their services.
A disrupted cluster, an unfinished defense
The attribution remains narrower than the overall activity. OpenAI says it is unclear whether all observed operators came from a single actor. Its assessment links a core cluster to individuals associated with Moonshot AI; it does not identify every operator as part of that group.
OpenAI frames unauthorized reasoning extraction as a safety and national-security risk. It argues that extracted reasoning could train another model without preserving the safeguards attached to the original model’s answers. At scale, it says, that could transfer advanced capabilities without the same investment in safety, with greater concern in domains where capabilities have both beneficial and harmful uses.
Before publishing, OpenAI deployed mitigations and shared findings with researchers and industry partners. It also passed information through the Frontier Model Forum and government channels. Investigation and mitigation continue. The next work includes extending protections to partner-hosted deployments and strengthening defenses against attacks carried through tool outputs—not just ordinary visible text.
Sources
openai.comDisrupting a coordinated model-distillation campaign
Reader comments
Newest comments first. Replies stay oldest first.