Anthropic Says Chinese AI Labs Used Millions of Claude Chats to Train AI Models

The global AI race just ran headfirst into a massive intellectual property and security breach. According to a detailed threat intelligence report from Anthropic, several prominent Chinese AI labs including Alibaba, Moonshot AI, and DeepSeek systematically harvested Claude’s outputs at an industrial scale to train and refine their own frontier models.

This was not casual API scraping. Anthropic disrupted multi-million-query campaigns designed to bypass terms of service and extract proprietary reasoning capabilities through what the industry calls “illicit distillation.”

Between December 2025 and August 2026, the volume reached staggering heights: Alibaba alone allegedly routed over 151 million exchanges through more than 3,500 fraudulent accounts to train its Qwen models, while DeepSeek generated over 12 million distillation queries in a two-week window in July.

The Mechanics of Illicit Model Distillation and Proxy Harvesting

To understand how this operation worked, you have to look at the economics of training large language models.

Pre-training a frontier-class model from scratch requires tens of millions of dollars in compute, specialized clusters, and massive curated datasets.

Model distillation, however, is a known shortcut: you feed prompts to a superior “teacher” model (in this case, Claude Opus) and use its structured answers, reasoning chains, and code outputs as synthetic training data for a “student” model.

What Anthropic uncovered went far beyond synthetic prompt generation. According to the report, Moonshot AI operated a covert relay mechanism.

When end users submitted queries to Moonshot’s Kimi interface, the system silently forwarded those requests to Claude via a distributed network of more than 5,300 proxy accounts routed through Singapore and Japan.

Moonshot then returned Claude’s answers back to unsuspecting Kimi users while intercepting and logging Claude’s full reasoning transcripts behind the scenes.

This enabled the lab to train its proprietary models on high-grade chain-of-thought data while simultaneously offloading expensive frontier inference costs onto Anthropic’s infrastructure.

Alibaba similarly automated queries at a peak volume of nearly three million hits per day, directly funneling Claude’s outputs into reinforcement learning pipelines and architecture development.

The Real Cost: Data Privacy and the Frontier IP Standoff

The biggest concern here is not just competitive poaching; it is the massive, unvetted pipeline of sensitive user data.

Because Moonshot and DeepSeek reportedly forwarded live customer prompts directly into Anthropic’s ecosystem without user consent, private queries ranging from corporate internal data to state-affiliated research were exposed to third-party infrastructure.

For enterprise IT teams and regulators, this shatters basic assumptions about data residency and vendor trust. If a user inputs proprietary financial projections or proprietary code into one local service, only for that request to be quietly handed off to another company’s model across borders, every conventional compliance framework collapses.

The report marks a major shift in how frontier AI labs will handle platform defense. Traditional security controls such as rate limits, account verification, and heuristic anomaly detection are no longer sufficient against well-funded research teams running distributed proxy botnets.

Moving forward, providers will likely tighten API access through rigorous cryptographic attestation, hardware-level fingerprinting, and aggressive behavioral filtering to detect synthetic query patterns.

As frontier models become more expensive to produce and harder to differentiate, the temptation to siphon distilled reasoning will only intensify.

This episode proves that the most valuable commodity in AI is no longer raw compute it is the hard-earned reasoning data of the leading models, and defending it is becoming an active cybersecurity front.

Source: CNBC, "Chinese AI Labs Moonshot, DeepSeek, Alibaba and Anthropic..."
Pradeepa Sakthivel
Pradeepa Sakthivel

Pradeepa is an AI Enthusiast and Technology Journalist covering AI News, AI Tools, Product Reviews, Industry Updates, and other developments in the rapidly evolving world of artificial intelligence.

Articles: 325