Behind the Blog: Rare Books and Baseball Brain
404 Media Samantha Cole
404 Media's story on AI firms buying used books to scan for training data went viral on X, but a copy-paste engagement account got the credit. Journalist Emanuel says it also twisted what the piece actually said.
Based on reporting by 404 Media, Samantha Cole — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Two weeks ago, 404 Media's Emanuel Maiberg published a story about AI companies quietly buying up printed books so they can scan the pages and feed the text into training pipelines. It's a small, weird corner of the AI supply chain: instead of licensing ebooks or scraping the web, some outfits are apparently going old school, snapping up physical copies to digitize on their own terms. The story is exactly the kind of granular, reporting-driven piece 404 Media built its name on.
But the way it spread tells its own story about how information moves online in 2024. Rather than the article itself going viral, an engagement-farming account on X lifted the gist, repackaged it, and pushed it out without a proper link back to 404 Media. The post took off. Millions of eyeballs, presumably, landed on a garbled version of the reporting instead of the source. Maiberg says the summary also misrepresented what he actually found, which is arguably worse than just losing the traffic.
This is the modern content-laundering problem in miniature. Independent outlets like 404 Media survive on subscriptions and direct traffic, not ad impressions, so having a story stripped of context and repurposed for someone else's engagement metrics is a direct hit to the business model. And it muddies the actual news: readers walk away with a distorted picture of how AI companies are sourcing training data, based on a paraphrase of a paraphrase rather than the reporting itself.
The underlying story deserves better than that treatment. Buying physical books to scan for AI training says something real about how much appetite there still is for clean, well-formatted, copyright-adjacent text that isn't already floating around the open web. That's a story about scale, cost, and how far AI labs will go to find fresh data. It got flattened into a viral tweet, and the outlet that did the actual legwork got left holding an SEO problem instead of the credit.
My take — AI-written commentary, not fact-checked reporting
Engagement farming accounts ripping off original reporting isn't new, but AI's data hunger has made the irony sharper: a story about AI companies extracting value from other people's work got extracted for value itself before most readers ever saw the source. Small outlets doing real reporting on AI's messy supply chains deserve links, not screenshots.
Read more about this at: 404 Media