Instead, the campaign relied on manipulating model interactions. One technique involved copying encrypted reasoning from one conversation and then asking a model in a separate conversation to decrypt and transcribe the hidden content. OpenAI said the method could reproduce protected reasoning in a form visible to the requester without compromising the underlying encryption system.
Activity was initially limited after appearing on July 1, according to OpenAI. The company later observed a sharp increase on July 24 and 25, when 16,000 requests matching the extraction pattern came from more than 4,000 users.
A broader investigation linked related prompt activity to a cluster of more than 15,000 users. OpenAI said it had shut down the campaign by July 28.
The company stopped short of saying every participant was working for the same organization. It said only that a central group of operators was associated with Moonshot AI.
“This activity is consistent with adversarial distillation: the systematic and unauthorized use of one model’s outputs or reasoning to help train, reproduce, or improve another model,” OpenAI wrote.
OpenAI has increasingly limited users’ access to full model reasoning, instead returning final answers while keeping internal reasoning protected. The company argues that extracting those traces at scale could make it easier for another developer to reproduce capabilities without the same investment in training or safeguards.
“Our concern is about violation of our terms of service, not open models or legitimate distillation,” Caroline Zier, who leads strategic national security policy initiatives at OpenAI, told Bloomberg.
OpenAI said it responded by banning or restricting accounts, tightening signup and infrastructure controls, expanding monitoring and adding protections around hidden reasoning. It also closed a pathway that allowed someone already holding another user’s encrypted reasoning to replay it and recover its contents.
The company said it added checks designed to detect streamed output that could expose protected reasoning. When related activity moved through third-party services, OpenAI said it worked with those providers to identify and block the accounts involved.
Independent security researchers also reported related vulnerabilities involving cross-model interactions and conversation compaction. OpenAI said it confirmed that those attack paths were real and used the findings to improve its defenses.
The disclosure adds to an expanding dispute over alleged distillation involving Chinese AI developers. Anthropic accused Moonshot AI in a report last month of routing large numbers of requests to Claude models and using the responses in ways that could support Kimi training.
U.S. agencies have also raised similar allegations against Chinese companies, while China has rejected those claims. Moonshot AI did not immediately respond to a request for comment.
The issue has also drawn political attention in Washington. OpenAI said it shared information about the campaign with other AI developers through the Frontier Model Forum and through government information-sharing channels.
The company argues that adversarial distillation creates both safety and national security concerns because extracted reasoning could be used to train another model without preserving the safeguards attached to the original system. OpenAI said the risk becomes more significant as models gain stronger capabilities in dual-use areas.
OpenAI, Anthropic and Google have begun coordinating on defenses against unauthorized distillation. U.S. lawmakers have also proposed legislation that would provide a limited antitrust exemption for companies sharing information about AI security threats, including model extraction.
OpenAI said it expects future attempts to become more sophisticated. Its response will focus on stronger technical protections, broader detection and enforcement, and more information sharing with industry and government partners.
The company also said the problem extends beyond its own infrastructure. Partner-hosted models and third-party services may require similar protections, particularly when model outputs or reasoning artifacts can be replayed across systems.
This analysis is based on reporting from BigGo Finance & OpenAI.
Image courtesy of Digital Watch Observatory.
This article was generated with AI assistance and reviewed for accuracy and quality.