TLDRocket
Sign in

đź”® Copy that: The curious case of AI distillation #594

Exponential View Azeem Azhar â—Ź Covered by 75 sources

Chinese AI labs may be copying US model outputs, and no one's sure if that's actually illegal. Anthropic says DeepSeek, Moonshot and MiniMax scraped 16 million Claude chats using 24,000 fake accounts.

Based on reporting by Exponential View, Azeem Azhar — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

There's a neat historical parallel buried in this week's distillation debate. Back in 1789, Samuel Slater memorized Richard Arkwright's cotton-spinning system, disguised himself as a farm laborer, and smuggled that knowledge out of Britain in his head, since Britain had banned exporting textile machinery and even banned skilled workers from leaving the country. He built a fortune in America worth roughly a billion dollars in today's terms. He knew he was breaking the law. That's the key difference from what's happening in AI right now.

Distillation itself is nothing new, a decades-old technique where a larger model trains a smaller one. Nobody's arguing with a lab distilling its own model internally. The controversy is about external distillation, allegedly without permission, and it's the accusation currently aimed at Kimi and other Chinese labs. A former OpenAI staffer named Almeida notes that this used to be easier, back when models happily exposed their reasoning traces. Now labs are mostly limited to what he calls behavior parroting, learning from final answers alone, a weaker technique but apparently still useful for bootstrapping a new model. Anthropic claims DeepSeek, Moonshot and MiniMax pulled more than 16 million Claude chats through 24,000 fake accounts. Trump's science chief, Michael Kratsios, says he has evidence of Moonshot running distillation attacks too. And it's not one-directional: Anastasios Angelopoulos of the benchmarking company Arena argues Kimi K3 is beating some top US models in ways distillation alone can't explain, and predicts American labs will soon start distilling Chinese intelligence right back.

Here's the uncomfortable bit nobody wants to say out loud: it isn't clear any of this is actually illegal. There's no legal precedent that model outputs count as intellectual property, and the US Copyright Office's 2023 statement backs that up, saying that when an AI determines the expressive content of its own output, that output isn't the product of human authorship and therefore isn't copyrightable. So the technology has simply outrun the law. And if labs did somehow claim IP over their models' outputs, they'd effectively own a slice of every future economic activity their users ever undertake, which is a strange position for labs to want, given they built their own systems by training on other people's books and essays.

Bahrad Sokhansanj offers a more grounded path: if distillation needs addressing, do it narrowly, tied to a specific harm, whether that's national security or something else, rather than writing sweeping rules that just hand the incumbent labs more power to shape the regulatory field in their favor. There's precedent for that worry too. Once Slater was a wealthy, established American industrialist, he lobbied for protectionist tariffs to keep foreign competitors out. Back in his hometown, they started calling him Slater the Traitor.

My take — AI-written commentary, not fact-checked reporting

Nobody actually wants a clean answer here, because a clean answer cuts against whoever's currently ahead. Labs built their empires by hoovering up the entire internet's worth of human writing without asking, so crying foul the moment someone does the same to their model outputs looks more like a business complaint than a principle. The real fight isn't about right and wrong, it's about who gets to write the rules before the other side catches up.

Read more about this at: Exponential View

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.