Detecting and Preventing Distillation Attacks
Key point
Anthropic confirmed that Chinese AI companies carried out large-scale distillation attacks to illicitly extract Claude's capabilities.
Details
Anthropic confirmed that three AI labs—DeepSeek, Moonshot, and MiniMax—ran an industrial-scale campaign to illicitly extract Claude's capabilities. They mobilized roughly 24,000 fraudulent accounts to generate over 16 million conversations, violating terms of service and regional access restrictions.
The distillation technique they used trains a lower-performing model using the outputs of a powerful model. While this can also serve as a legitimate method for efficient model development, it can be abused by competitors to steal another company's core capabilities without the massive cost and time investment normally required.
Such attacks pose serious national security threats. Illicitly extracted models are likely to be deployed with the safety guardrails built by Anthropic removed, which could be exploited for dangerous activities such as developing biochemical weapons or carrying out cyberattacks.
In particular, this behavior by Chinese companies appears to be an attempt to neutralize U.S. export control measures. DeepSeek was found to have intensively trained on reasoning ability, reward modeling, and generating censorship-evasive responses through more than 150,000 exchanges.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.