Who’s Afraid of Chinese Models?
Ben Thompson proposes US legislation to clarify that training data collection is fair use and to ban terms that forbid distillation, countering Chinese open-source model competition and revealing the hypocrisy in AI policies.
- Ben Thompson points out the hypocrisy of US AI labs banning distillation while training on unlicensed data themselves.
- He proposes legislation: affirm data collection as fair use, and ban terms of service that forbid distillation, unleashing innovation.
- Chinese AI models like Qwen iterate rapidly through open source and distillation, potentially influencing US policy discussions.
- Distillation is essentially knowledge transfer; technical barriers should not become legal ones, and it may reshape AI competition.
Recently, several major US AI labs have added “no distillation” clauses to their model usage terms, aiming to prevent other companies from using their API outputs to train competing models. Tech analyst Ben Thompson pointedly noted the hypocrisy: these same companies scraped vast amounts of public internet data for training without obtaining licenses from all copyright holders, yet now they forbid others from learning from their models in a similar way.
Simon Willison referenced Ben’s proposal in his blog: the US should pass a law that both clarifies collecting data for model training as “fair use” and bars companies from using terms of service to prohibit distillation. The rationale is simple – distillation is essentially just querying an API, and it’s nearly impossible to stop technically. Instead of futile restrictions, the US should embrace openness, allowing knowledge to flow freely and sparking more innovation.
Even more intriguing is Alibaba’s recent release of Qwen 3.8 Max as open weights, a reversal from the closed-source 3.7 Max. Simon speculated that this shift may have been influenced by Xi Jinping’s recent speech emphasizing “encouraging open source, openness, collaboration, and sharing.” This highlights how policy in China directly drives AI open-source ecosystems, potentially accelerating US policy rethinking as well.
Trend Insight
This episode reveals that global AI competition is moving from a model performance arms race to a battle over rules and ecosystems. The rise of open-source models breaks closed-source monopolies, and distillation allows model knowledge to be extracted and reused. Closed-source companies trying to maintain advantages through legal and technical barriers may be swimming against the tide. If the US passes such legislation, it would dramatically lower the barriers to model innovation, enabling small teams to build on top of top-tier models, reminiscent of the early internet’s openness.
Practical Value
For AI practitioners, it’s time to reframe distillation: it’s not theft, but a legitimate and potentially policy-encouraged means of knowledge transfer. If US law changes, global developers will have easier access to advanced model capabilities, though they must still navigate compliance, especially with cross-border data flows. Meanwhile, the success of Chinese open-source models shows that openness is key to competitiveness; companies shouldn’t just focus on closed-source moats.
Counterintuitive Angle
Many worry that distillation will kill the business of large model companies. But the value of a model lies not just in its weights, but also in data flywheels, user ecosystems, and service quality. Allowing distillation could actually increase traffic and revenue for base model companies via their APIs, much like search engines allow crawling but profit from ads. Distillation isn’t a zero-sum game; it’s a way to expand the innovation pool.
While China says “encourage open source” and the US considers “distillation freedom,” these seemingly coincidental policy moves actually signal a broader shift from wild growth to rule-making in global AI. For everyday developers, a more open and vibrant era of AI innovation may be arriving faster than we think.
Analysis by BitByAI · Read original