Quick Takeaways
- Chinese AI labs are conducting large-scale illicit distillation attacks by rerouting user requests through proxy accounts, stealing sensitive model capabilities and training data from Anthropic’s Claude AI.
- These illicit campaigns involve creating thousands of fake accounts using stolen credentials and harvesting user conversations, including sensitive and proprietary information, for unauthorized model training.
- Increased use of proxy networks and data resale markets enables persistent circumvention of access restrictions, threatening intellectual property, user privacy, and the integrity of AI models.
Threat, Attack Techniques, and Targets
Anthropic has identified and disrupted multiple large-scale distillation attacks against its AI model, Claude. These attacks are mainly carried out by seven China-based AI labs, such as Alibaba, Moonshot, DeepSeek, Z.ai, and MiniMax. The main goal is to extract and copy the capabilities of Claude without permission. These labs use illegal methods to access the models and steal data.
The attack methods involve creating many fake accounts using stolen credit cards, fake identities, and illegal API keys. The labs route their requests through proxy services, making it hard to trace the activity. They also reroute user requests from their own models to Claude without user knowledge. These labs then save the conversations and responses from Claude to train their own models. Sometimes they even buy transcripts of user exchanges from third-party resellers. They focus on advanced tasks like reasoning, coding, and data analysis.
The targets are the Claude AI models and their user data. The attacks aim to copy Claude’s capabilities and harvest sensitive conversations for training other malicious models or conduct illegal research.
Impact, Security Implications, and Remediation Guidance
The impact of these attacks includes the potential for stolen AI capabilities, misuse of user data, and violation of user privacy. The labs’ ability to copy advanced reasoning skills increases the risk of malicious and competitive AI use. This also undermines trust in AI systems and raises concerns about data security.
These attacks show that even with defenses, malicious labs can find ways to bypass security measures. Anthropic has responded by updating its models to include internal reasoning summaries and encrypting sensitive parts of its system. It also bans accounts from unsupported regions and uses verification steps to prevent unauthorized access.
For organizations using AI provided by vendors like Anthropic, it is important to stay updated on security features. If you suspect your data has been compromised or you need detailed guidance, contact the AI vendor or cybersecurity authorities. Remediation strategies should be obtained directly from the relevant vendor or security agency to ensure proper protection.
Discover More Technology Insights
Learn how the Internet of Things (IoT) is transforming everyday life.
Discover archived knowledge and digital history on the Internet Archive.
ThreatIntel-V1
