Is it legal to train AI models on copyrighted books? It’s complicated
TechCrunch Amanda Silberling
AI companies are getting sued over training on books, but the law isn’t settled. One judge even called Anthropic’s training lawful while still hitting it with a huge payout.
Based on reporting by TechCrunch, Amanda Silberling — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
If you’ve ever wondered whether the engines behind ChatGPT, Gemini, Claude and the rest are built on a legal mess, the answer is yes. They’re trained on enormous piles of books, articles and papers, much of it scraped without authors knowing. That sounds like an easy copyright violation. It isn’t.
One reason is that copyright law is still running on rules from 1976. Judges are being asked to apply old statutes to systems that can swallow trillions of words, and the results are messy. Cathy Gellis, an intellectual property attorney, says the whole area is loaded with tension and not nearly as simple as people want it to be.
The Anthropic case shows how weird this can get. Judge William Alsup ordered the company to pay $1.5 billion to writers, but he also said the training itself was lawful. The part he punished was the use of books taken from illegal shadow libraries. He compared an LLM’s reading of text to how a writer studies literature, which is not exactly a sentence that leaves everyone happy.
Gellis sees that ruling as a win for AI companies more than authors. Her point is blunt: copyright law is about copying, not just using or reading. And if Anthropic is headed toward about $200 billion in annual revenue by 2028, $1.5 billion starts to look less like a deterrent and more like a line item.
The bigger fight is fair use, and courts keep coming back to the same question: is the use transformative, or is it just a shortcut to a rival product? In one case, Thomson Reuters won when Judge Stephanos Bibas said Ross Intelligence’s use of Reuters material was not transformative because it was building a direct competitor. But that logic has not yet stuck when authors argue that chatbots are competing with them by generating synthetic books. For now, the lawsuits are still doing the law’s job for it, one awkward ruling at a time.
My take — AI-written commentary, not fact-checked reporting
The real scandal here isn’t that the law is confused; it’s that the AI industry treated “we found it online” like a moral warranty. Courts are only now doing the basic accounting. The rest of the hype machine can wait its turn in line behind copyright.
Read more about this at: TechCrunch