Concerns Over Chinese AI Model Kimi K3 Raise Questions About Technology Transfer
Recent statements from Michael Kratsios, a science advisor to the White House, have sparked significant concern regarding the development of the Kimi K3, a large language model (LLM) created by the Chinese company Moonshot. Kratsios alleged that the Kimi K3 was built by copying Anthropic’s Fable LLM and that Moonshot utilized advanced chips that are not authorized for export to China. His remarks highlight ongoing discussions about potential bans on Chinese open-weight models, a topic generating considerable debate within the artificial intelligence sector.
In a tweet, Kratsios emphasized that “large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable.” However, Moonshot has not responded to inquiries regarding its training methodologies, and Kratsios has not provided additional details to substantiate his claims.
Treasury Secretary Scott Bessent echoed Kratsios’ sentiments, stating that “we are finding watermarks of our U.S. large language models on many of the Chinese models, and that’s unacceptable.” Yet, the nature of these watermarks remains unclear, and the Treasury Department has not commented further on the issue.
Experts in the field have expressed skepticism regarding the effectiveness of the distillation process, which involves querying an LLM to replicate its capabilities. Braden Hancock, a researcher at the Laude Institute and co-founder of Snorkel AI, pointed out that achieving the advanced functionalities of Kimi K3 in such a short timeframe—especially given that Fable was only made publicly available in early July—would be implausible through distillation alone.
Nathan Lambert, an AI researcher at the Allen Institute for AI, noted that the impact of distillation appears to be diminishing as Chinese models evolve. He suggested that replicating capabilities akin to those of Fable would likely necessitate more sophisticated reinforcement learning techniques, which involve complex infrastructures and significant resources.
The distillation process itself requires systematic querying of a target model to generate data for post-training. This can include prompting the model to express its reasoning processes or utilizing its responses for supervised fine-tuning (SFT). Lambert remarked that while SFT can enhance a model’s performance, its advantages are becoming less significant as models grow increasingly complex.
Concerns regarding Moonshot’s practices are not isolated. Earlier this year, Anthropic accused Moonshot and other companies of systematically distilling its models. The company claimed to have identified millions of interactions indicative of deliberate capability extraction rather than normal usage.
Interestingly, the practice of distillation is not limited to Chinese firms. Elon Musk testified that his company, SpaceXAI, employed a similar approach to develop its AI model, Grok, suggesting that such practices are prevalent across the industry.
Hancock emphasized the technical capabilities of Chinese AI teams, noting that many are composed of highly skilled researchers. He cautioned against underestimating their expertise, stating that if American models were to stagnate, Chinese progress would likely continue.
Another critical aspect of Kratsios’ allegations involves the procurement of advanced Nvidia chips, specifically Grace Blackwell 300s, which are banned from export to China. Sam Bresnick, a research fellow at Georgetown’s Center for Security and Emerging Technology, highlighted the existence of a black market for these chips, raising concerns about accountability in data centers that may be facilitating unrestricted access to cutting-edge technology.
President Biden’s Department of Commerce has proposed federal regulations to improve transparency in data centers, but progress on these initiatives has been slow. Current regulations require exporters of advanced chips to ensure their intended use aligns with approved purposes.
As discussions around AI technology transfer continue, the implications of these developments could have lasting effects on both national security and the global AI landscape.

