TLDRocket
Sign in

Experts say exploiting Anthropic’s Fable isn’t how Kimi K3 got so good

TechCrunch Tim Fernholz Covered by 75 sources

White House officials claim China's Kimi K3 was built by copying Anthropic's Claude model using banned Nvidia chips. AI researchers say the timeline just doesn't add up for straight-up distillation.

Based on reporting by TechCrunch, Tim Fernholz — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Michael Kratsios, the White House's top science advisor, dropped a pretty serious accusation this week: Moonshot AI copied Anthropic's Fable model to build Kimi K3, and did it using Nvidia GB300 chips that are supposed to be off-limits to China. Treasury Secretary Scott Bessent piled on, claiming officials are finding "watermarks" of U.S. models scattered across Chinese systems. Neither Kratsios nor the Treasury Department has explained what evidence backs this up, and Moonshot hasn't said a word about its training process.

Here's the problem researchers keep bumping into: the math doesn't work. Fable only became publicly available on July 1st. Braden Hancock, a researcher at the Laude Institute, points out that distilling enough data from a model, training on it, and shipping a competitor in roughly two weeks just isn't realistic. "There's just not even frankly time," he told TechCrunch. Nathan Lambert at the Allen Institute for AI made a related point on a podcast this week — if simple distillation explained models like K3 or GLM, every lab chasing the frontier would have already caught up by copying data. They haven't.

The distinction that keeps getting flattened in all this is between old-school supervised fine-tuning, where a smaller model mimics prompts and responses from a bigger one, and the reinforcement learning approaches that actually produce frontier-level reasoning. SFT is cheap and explains why some knockoff models occasionally introduce themselves as Claude. But matching genuine capability gains, according to Lambert, increasingly requires RL setups with tens of millions of grading agents — infrastructure that would be brutally expensive and slow if you tried to run it through someone else's API.

None of this means Chinese labs are innocent of anything. Anthropic itself accused Moonshot, DeepSeek, and MiniMax back in early 2025 of systematically querying its models in patterns that looked like deliberate extraction rather than normal use. And distillation broadly is just how the industry works now — Elon Musk testified that his own team used OpenAI's outputs to help build Grok. The chip question is murkier still: Sam Bresnick at Georgetown's Center for Security and Emerging Technology notes a black market for banned Nvidia hardware clearly exists, and a Supermicro executive was indicted in May for smuggling chips into China. Commerce Department rules requiring data centers to actually know their customers were proposed under Biden but never finalized, and nothing's moved on that front under Trump.

Hancock's blunter point may matter more than the accusation itself: underestimating Chinese AI teams has been a recurring American mistake. One of Moonshot's founders did a PhD at CMU. These aren't hobbyists riding on stolen homework — they're trained researchers building real systems, chip restrictions or not.

My take — AI-written commentary, not fact-checked reporting

This smells like officials reaching for a simple villain story because the real one — that Chinese labs have legitimately competent researchers and RL techniques that don't require anyone's secret sauce — is less politically satisfying. Export controls on chips are a real and enforceable lever; vague talk of "watermarks" with no evidence is not policy, it's vibes. If the U.S. wants to slow China down, tightening actual chip supply chains and enforcing know-your-customer rules will do far more than tweeting accusations nobody can verify.

Read more about this at: TechCrunch

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.