Adversarial distillation is the kind of attack that doesn’t make headlines the same way a data breach does, but it might be more consequential. OpenAI reported that it identified and shut down a coordinated campaign designed to extract the internal reasoning of its models, with the earliest activity traced back to the first week of July. The core finding: a cluster of operators, some linked to Moonshot AI, the Chinese company behind the Kimi chatbot, ran a scaled, systematic operation to pull protected reasoning out of OpenAI’s models without authorization.
What they were actually after
Protected reasoning is the model’s internal chain of thought, the working-out process before a final answer is produced. It can contain information the model deliberately withholds from its visible output. If you can extract that reasoning and use it to train another model, you’re essentially copying capabilities, potentially without preserving the safety guardrails that were built into the original. That’s the threat here. It’s not about stealing user data or breaking encryption. It’s about copying intelligence.
The attack method was novel. Operators copied encrypted reasoning from one conversation and asked the model in a separate conversation to decrypt and transcribe it. No database was compromised. No encryption was broken. The manipulation worked through the model interaction itself, which makes it harder to detect and harder to attribute.
The scale of the campaign
Activity started quietly on July 1. Then on July 24 and 25, volume spiked sharply: 16,000 requests using a relevant extraction pattern, coming from over 4,000 users. Further investigation found related prompt-pattern activity across a cluster of more than 15,000 users. OpenAI says it fully disrupted the campaign by July 28.
Independent security researchers also flagged related vulnerabilities through responsible disclosure, specifically cross-model and conversation-compaction attack paths. OpenAI confirmed those were real and said the outside research helped accelerate its mitigations.
Why Moonshot AI’s involvement matters
OpenAI stopped short of saying all the activity came from a single actor. But it directly attributes a core cluster to individuals associated with Moonshot AI. That’s a significant public attribution. Moonshot is a well-funded Chinese AI startup competing in the same tier as other frontier model developers. The implication is that competitive pressure is pushing some actors toward extracting capabilities rather than building them independently.
This isn’t a problem unique to OpenAI. The same techniques could work against any model that exposes reasoning artifacts. OpenAI shared its findings through the Frontier Model Forum and government information-sharing channels, which is the right call. But it also means every frontier lab needs to be treating adversarial distillation as a live threat, not a theoretical one.
What OpenAI changed and what’s still open
The response covered several fronts:
- Banned or restricted fraudulent accounts and strengthened signup controls
- Closed the pathway that allowed someone with encrypted reasoning to replay and recover its contents
- Added detection for streamed output that might expose reasoning
- Worked with third-party services to identify and shut down accounts routing through them
- Shared findings with industry partners and government channels
But OpenAI is explicit that the work isn’t done. Partner-hosted deployments need the same protections as first-party services. Tool-output attacks require defenses that go beyond scanning visible text. Classifier coverage, model refusals, and controls across cloud partners are all still being improved.
The broader picture here is that as frontier models get more capable, the incentive to steal rather than build grows. Distillation attacks are cheaper than training from scratch and harder to detect than a conventional breach. Expect them to get more sophisticated. The industry’s collective response, how quickly labs share threat intelligence and how seriously they treat this attack class, will shape whether distillation becomes a persistent drain on AI research investment or something the field learns to contain.



