For years, model distillation operated in the shadows of API logs and proxy traffic – a cat-and-mouse game of rate limits and pattern matching. That changed in September 2026, when Anthropic’s threat intelligence team published Case GTG-16008, detailing a systematic campaign by Xiaomi to harvest data from Claude to train its own models. This was not a rogue developer scraping for a side project. It was an industrial-scale operation that reveals distillation as the primary development strategy for a significant portion of the Chinese AI ecosystem.
The Xiaomi campaign was precise and aggressive. Over a 20-day window in March and April 2026, the company funneled more than 400,000 requests to Claude through 1,500 distinct accounts, using proxy services to mask the traffic’s origin. Xiaomi’s engineers used coding harnesses like OpenClaw and OpenCode to replay actual user conversations and coding sessions from their own MiMo models through Claude’s interface.
The objective was not to serve better answers to Xiaomi’s users. Anthropic’s investigation found that Xiaomi used Claude to reconstruct developer environments from exchange transcripts, clean up multi-turn conversations into structured training pairs, generate synthetic input-output exchanges that mimicked human-model interactions, and judge the quality of its own model’s outputs. The timing is revealing: Xiaomi launched its MiMo-V2-Pro with a free trial period, and the bulk distillation began as that trial concluded – suggesting the trial was used to generate the developer traffic that fed the harvesting pipeline.
The Xiaomi campaign sits inside a much larger pattern. Anthropic has identified seven PRC-based labs conducting industrial-scale distillation, with a combined total exceeding 200 million exchanges. Alibaba alone accounted for approximately 151 million exchanges across nearly 5,000 fraudulent accounts, peaking at 3 million per day. Moonshot relayed almost 300,000 customer requests to Claude in a single ten-day window, silently serving Claude’s responses to users who believed they were using Kimi. DeepSeek built a similar relay pipeline, routing requests from users who thought they were interacting with DeepSeek’s own models. The CISA Advisory AA26-251A, issued jointly by the NSA, CISA, and FBI on September 8, 2026, names six Chinese AI companies conducting these campaigns since late 2024 – confirming this is a coordinated, state-level effort, not isolated corporate misconduct.
Beyond the intellectual property theft, there is a privacy dimension that implicates both the labs and the proxy ecosystem that enables them. In their rush to distill Claude, these labs relayed sensitive data from hundreds of their own users – names, contact information, proprietary corporate data, even live credentials – through third-party routing services commonly accessed by users in the United States and Europe. The data crossed at least a dozen languages. None of the users whose data was exposed consented to, or were likely aware of, the relay. Moonshot’s pipeline included a PLA-affiliated user’s CCTV surveillance data from Chengdu. DeepSeek’s relay exposed live credentials for a Russian government database. These are not abstract risks – they are the operational cost of treating user data as raw material for distillation.
The connection between these campaigns and the rapid ascent of open-weight models from the same labs is the question the industry has been avoiding. Xiaomi released MiMo-V2.6 – an open-weights model that ties Grok 4.7 on the Intelligence Index – the same week Anthropic published its threat report. The model’s 1.02 trillion total parameters and frontier-class benchmarks arrived after a training period that, according to Anthropic’s findings, included data harvested from Claude. The timing does not prove that distillation drove the performance gains, but it narrows the gap between the harvesting window and the capability jump in a way that makes the provenance question unavoidable.
Anthropic’s response has been technical and structural. With the release of Claude Opus 5.5 – the same day its threat report went live – the company introduced Fable 5.1 safeguards that target distillation directly. The model now summarizes its internal reasoning before responding, making stolen transcripts less useful for training. A new “preserved thinking” feature prevents new API accounts from altering the system prompt or tools that precede Claude’s reasoning – closing the cross-session replay attacks that Moonshot and DeepSeek used to extract chain-of-thought transcripts. The thinking process cannot be disabled. These are not incremental adjustments; they are architectural changes designed to break the feedback loops that industrial-scale distillation depends on.
The timing compresses the story. In the same 48-hour window: Anthropic published its threat report and shipped Opus 5.5 with anti-distillation safeguards; OpenAI released GPT-6 Sol and Luna at a 50% price cut; Xiaomi’s MiMo-V2.6 was already live at frontier-class performance; and the Bessent management-responsibility doctrine was being tested in court via the BC AG v. OpenAI lawsuit. The pricing war, the legal accountability arc, and the distillation enforcement action are converging into a single structural shift: frontier labs are no longer just building models – they are defending the intelligence layer of the global AI economy against systematic extraction by state-backed competitors.
For builders evaluating the open-weight ecosystem, the implications are direct. If frontier-class performance from labs identified in the CISA advisory is built in part on unauthorized distillation, the sustainability and licensing integrity of those models come into question. Xiaomi has not publicly responded to Anthropic’s findings. The antitrust lawsuit filed September 18 alleged that US labs coordinated to slow the frontier. The distillation data suggests that while US labs were debating the pace of release, their Chinese competitors were extracting the capabilities they had already built – and shipping them as open weights that undercut the pricing structure the coordinated slowdown was meant to protect.
