OpenAI disrupts reasoning extraction campaign linked to Moonshot AI
OpenAI blocked a coordinated effort to extract protected reasoning traces from its models, attributing the activity to associates of Chinese firm Moonshot AI.
Bài viết này chỉ có sẵn bằng tiếng Anh.
OpenAI identified and dismantled a large-scale campaign designed to illicitly extract protected reasoning data from its artificial intelligence models. The company attributed the core of this activity to individuals associated with Moonshot AI, a Beijing-based competitor, and fully disrupted the operation in late July 2026.
What happened
The campaign began on July 1, 2026, starting with low-volume attempts that gradually escalated. By July 24 and 25, the activity spiked significantly, with over 4,000 users generating approximately 16,000 attempted requests using specific extraction patterns. Further investigation revealed that related prompt-pattern activity involved more than 15,000 users across the platform. OpenAI fully disrupted the campaign on July 28, 2026, banning the fraudulent accounts involved.
OpenAI stated that the operators did not break encryption, compromise databases, or gain direct access to stored user conversations. Instead, they manipulated model interactions to reproduce protected reasoning in forms visible to the requester. This method violated the company's terms of service and was characterized as adversarial distillation, which involves the systematic and unauthorized use of one model's outputs to train or improve another model.
This incident follows previous accusations against Moonshot AI. Last month, Anthropic accused the company of stealthily relaying customer requests to its Claude model instead of processing them with its own Kimi model, allegedly retaining some exchanges to train its chain-of-thought capabilities. The current activity has been tracked under the moniker GTG-16002.
How it works
The attack leveraged an architectural vulnerability detailed in a study published in August 2026 by researchers from MATS Research, ELLIS Institute Tübingen, and Synk. The study found that encrypted reasoning traces were fully compatible and interchangeable across different sessions, users, and models within a provider's ecosystem. This technical flaw allowed attackers to develop a scalable decryption jailbreak.
By injecting an encrypted reasoning trace from a powerful model into a weaker, less safeguarded model from the same provider, attackers could force the weaker model to decode and output the trace in plaintext. This process bypassed the need to directly jailbreak the more capable model. The technique also enabled large-scale private data extraction and invisible prompt injections by embedding malicious payloads within encrypted blocks.
Key details
- The campaign ran from July 1 to July 28, 2026, with a major spike on July 24-25 involving 16,000 attempts from over 4,000 users.
- OpenAI attributed the core activity to individuals associated with Moonshot AI but did not cite specific technical evidence for this attribution.
- The attack method involved manipulating model interactions rather than breaking encryption or compromising databases.
- A related study revealed that encrypted reasoning traces were interchangeable across models, allowing weaker models to decode traces from stronger ones.
- OpenAI closed the pathway that allowed replaying encrypted reasoning to recover contents and added checks to detect exposed reasoning in streamed output.
- The activity is tracked as GTG-16002 and follows prior distillation accusations against Moonshot AI by Anthropic.
Why it matters
For software engineers and AI developers, this incident highlights the fragility of current safeguards around model reasoning. Protected reasoning traces offer insights into how a model processes tasks, and extracting them can reveal sensitive data or help reproduce capabilities without the original safety investments. As OpenAI noted, adversarial distillation poses safety and national security risks because extracted reasoning can be used to train other models without preserving the original safeguards.
The vulnerability also demonstrates that security is not just about preventing direct access but also about managing how models interact within an ecosystem. If encrypted data is interchangeable across models, a weakness in a less safeguarded model can compromise the entire system. This has implications for how providers design model architectures and handle encrypted data streams, especially as models gain capabilities in dual-use domains.
What you can do
- Monitor your API usage for unusual patterns or spikes in requests that may indicate coordinated extraction attempts.
- Review your terms of service and enforcement mechanisms to ensure they cover adversarial distillation and unauthorized data reproduction.
- Implement checks to detect and hold streamed output that might expose internal reasoning processes or protected data.
- Ensure that encrypted data traces are not interchangeable across different models or sessions within your ecosystem to prevent cross-model decoding attacks.
- Stay informed about emerging vulnerabilities in AI architectures, such as those highlighted in recent academic studies, to proactively adjust your security posture.
- Consider the safety implications of distillation when training or fine-tuning models, ensuring that safeguards are preserved in derived systems.



