TLDRocket
Sign in

Investigation finds distillation could threaten frontier model business profits

Business Insider Covered by 3 sources

AI labs are freaking out over "distillation" — training cheaper models on the outputs of pricier rivals. Turns out billions in R&D can be copied fast and cheap, wrecking the profit math behind frontier AI.

Based on reporting by Business Insider — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Distillation used to be a quiet, technical footnote in machine learning research. Back in 2015, Google researchers described it as a way to train a smaller model using the outputs of a company's own larger one — tidy, self-contained, uncontroversial. That era is over. Since ChatGPT's 2022 debut, distillation has become something closer to an industry-wide open secret, with companies quietly (or not so quietly) training on rivals' outputs to catch up fast and cheap.

The economics are what make this dangerous for the biggest players. Anthropic, OpenAI, and Google have poured huge sums into data, talent, and compute to build frontier systems they hoped to sell at premium prices. But if a rival can get most of the way there by learning from those models' outputs, the original investment starts to look like a subsidy for the competition. Anthropic's head of policy, Sarah Heck, put it plainly in a letter to US politicians: the practice inverts the logic that's supposed to reward American AI leadership. Elon Musk, sparring with OpenAI in court this year, was more blunt — "generally, AI companies distill other AI companies" — and when pressed on whether xAI had done it too, he answered, "partly."

China has become the flashpoint. Anthropic has accused Alibaba of running large-scale harvesting operations, including tens of thousands of fake Claude accounts built purely to scrape answers. Researchers have speculated that Z.ai's GLM-5.2 pulled from Claude and GPT systems, and AI stocks wobbled after that model's release. Former Meta product manager Xiaoyin Qu summed up the mood on X: labs burn billions building the newest model, only to watch free Chinese alternatives eat their margins. Google DeepMind's Yao Shunyu has argued that Chinese firms, often chip-constrained, have turned distillation — especially more advanced forms where multiple models refine each other's answers — into a genuine competitive edge.

But nobody agrees on where the line sits. Google itself paid Scale AI workers to generate and refine ChatGPT-style answers while racing to catch OpenAI, and plenty of researchers don't even count that as distillation. Meanwhile AI researcher Nathan Lambert warns that panic over the practice risks punishing small companies and academics who rely on cheap distillation techniques just to do research at all. Terms of service across the industry ban using outputs to build competitors, but enforcement is another matter — Oxford China Policy Lab researcher Zilan Qian describes a growing network of Chinese "transfer station" proxy services that let customers dodge account restrictions for a fraction of the official price, often using armies of fake accounts or paid identity-check workers in lower-income countries.

Anthropic's countermeasures illustrate the bind. The company blocked Chinese users, demanded overseas phone numbers and payment details, and asked some users for government ID and live selfies — then quietly degraded responses to AI-development questions before partially reversing course after developer backlash. Reports also surfaced that it dropped spyware meant to track Chinese users. Qian's blunt assessment: locking down access rarely stops determined users, it just raises the price of getting around it and creates a profitable market for whoever's willing to do it. Which, ironically, is pushing more developers toward exactly the cheap, distilled open-source models the frontier labs are trying to choke off.

My take — AI-written commentary, not fact-checked reporting

Watching billion-dollar labs cry foul over distillation is a bit rich when Google admittedly did its own version of the same thing to chase OpenAI. Everyone wants their own outputs protected and everyone else's fair game — that's not principle, that's just leverage running out. If frontier labs can't out-build cheaper distilled models on merit, tighter account verification and legal threats aren't going to save the business model.

Read more about this at: Business Insider

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.