TLDRocket
Sign in

Who’s Afraid of Chinese Models?

Simon Willison's Weblog Simon Willison Covered by 75 sources

Ben Thompson wants a US law to protect AI distillation and calls training data fair use. It could help US open models keep pace with China's, where Alibaba just went open weights on Qwen after a nudge from Xi Jinping.

Based on reporting by Simon Willison's Weblog, Simon Willison — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Ben Thompson has a proposal that cuts right through the AI industry's favorite double standard. Labs like OpenAI and Anthropic built their models by hoovering up unlicensed data from across the internet, calling it fair use. But when a rival — often a Chinese one — queries their API repeatedly to train a competing model, suddenly that's theft, and terms of service get rewritten to ban it. Thompson's fix: codify both halves of the deal into US law. Make training-data collection explicitly fair use, and outlaw contract clauses that forbid distillation, at least for American companies.

The practical argument is hard to dismiss. Distillation is, at bottom, just sending queries to an API and learning from the responses. You can't meaningfully stop that without shutting the API down entirely, and even then, someone will find a workaround. So instead of pretending distillation can be policed, Thompson says the US should lean into it — indemnify the labs against data-collection lawsuits while guaranteeing that whatever they learn from the web eventually becomes fuel for everyone else's models too. It's a trade: less legal risk for the giants, more downstream competition for the ecosystem.

There's a geopolitical wrinkle here that makes the timing interesting. Alibaba just flipped its stance and released Qwen 3.8 Max as open weights, a reversal from May's decision to keep Qwen 3.7 Max closed. Thompson floats the idea that this wasn't a purely commercial call — it may trace back to a Xi Jinping speech urging Chinese firms to treat this moment as a rare historic opening for open source collaboration and sharing. If Beijing is actively pushing its labs toward openness while releasing genuinely capable models for free, that changes the competitive math for every US company sitting on closed weights.

Which is really the crux of Thompson's pitch. If Chinese labs get state encouragement to open up their models and build goodwill, ecosystem lock-in, and developer mindshare in the process, American open-model efforts need something more than good intentions to compete. Clear legal ground for both training data and distillation wouldn't guarantee US labs win that race, but right now they're arguably tying one hand behind their back over a hypocrisy nobody in the industry wants to own up to.

My take — AI-written commentary, not fact-checked reporting

I like this proposal precisely because it forces AI labs to pick a lane instead of having it both ways — you don't get to call scraping the entire internet 'fair use' and then sue someone for doing the AI equivalent of reading your homework. And the Qwen reversal is the tell here: when a state-directed push toward openness starts outmaneuvering Silicon Valley's walled gardens, closed labs should be more worried about their own policy contradictions than about China.

Read more about this at: Simon Willison's Weblog

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.