Rohan Paul
@rohanpaul_ai
In a new allegation, the U.S. government has accused 6 Chinese AI firms of using large-scale distillation to copy American model capabilities.
The NSA, CISA and FBI published joint advisory AA26-251A, naming DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z .AI as running industrial-scale distillation campaigns against U.S. frontier models since at least late 2024.
The agencies say requests moved through a gray market of API proxies called transfer stations, bulk premium subscriptions shared across developer teams, and third-party aggregators that strip account metadata.
Moonshot AI allegedly distilled 17 U.S. models, including Anthropic's Claude Fable 5, to train Kimi-K3.
They claim DeepSeek's prompts pushed models to write out hidden chain-of-thought steps, which transfers reasoning method rather than finished answers.
So they say that DeepSeek's headline training cost only covers the compute it burned. The advisory argues that number is misleading because it leaves out what the training data actually cost, which DeepSeek allegedly obtained by distilling U.S. models rather than producing it through its own research. So the cheap-training story rests on data someone else paid to create.
The mitigation section asks U.S. labs to serve suspected distillers subtly degraded responses without telling them.
Its detection indicators are behavioral, covering sustained 24/7 usage, new accounts at immediate maximum throughput, and traffic optimized for cache hits.
Those patterns also describe an ordinary enterprise agent fleet, leaving each provider to decide which customers receive an undisclosed downgrade.