TLDRocket
Sign in

Suno snatched millions of songs from YouTube, Genius, and Deezer

The Verge Jess Weatherbed

A hack exposed Suno's actual training data: millions of songs scraped from YouTube Music, Deezer, Genius and more. It's the first real proof behind lawsuits claiming the AI music tool ripped copyrighted tracks.

Based on reporting by The Verge, Jess Weatherbed — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

For months Suno has stayed vague about what exactly went into training its AI music generator, leaning on a fair-use defense while keeping the actual dataset under wraps. A hacker going by "ellie.191" just made that vagueness a lot harder to maintain, handing 404 Media a trove of source code and scraping instructions from 2023 and 2024 that reads like a shopping list of the internet's music libraries.

The files reportedly show Suno pulling audio from YouTube Music, Deezer, Genius, Pond5, Jamendo, Freesound, and the International Music Score Library Project. One file recorded 2,013,545 YouTube Music clips consumed at the point it was last updated. Other datasets tallied hundreds of thousands of hours from YouTube Music, thousands of hours from Deezer, Genius, IMSLP, Jamendo and Pond5, and hundreds of hours pulled from Freesound and MuseScore lyrics. Leaked code also pointed to Suno using a company called Bright Data to scrape YouTube, and hunting specifically for a cappella tracks to isolate vocals. On top of the music, the company apparently tried to grab roughly a million hours of podcasts through a tool called PodcastIndex.

This lines up uncomfortably well with what the RIAA has already alleged in court. Suno has admitted outright, in filings tied to the RIAA's lawsuit, that it trains on copyrighted material scraped from the open internet, betting that fair use will cover it. The RIAA's amended complaint went further, accusing Suno of stream-ripping YouTube tracks by deliberately bypassing the platform's copyright protections — an allegation the leaked scraping instructions now seem to back up in granular detail.

Suno's response to 404 Media stuck to the same script it has used in court, saying its models were trained on "publicly available music files and related metadata" found on third-party sites. That framing does a lot of work, since scraping something doesn't automatically make it fair game, and a judge will eventually have to decide whether that argument holds.

The breach wasn't limited to training data. The hacker also pulled customer emails, phone numbers, and Stripe payment details. Suno told 404 Media it discovered the incident in November 2025, contained it quickly, and concluded the exposed material was outdated code and limited customer information — not enough, in its view, to legally require notifying affected users. Some of those customers, when contacted directly, said they'd never heard from Suno about any breach at all.

My take — AI-written commentary, not fact-checked reporting

A company that spends months insisting its training data is a mystery, only to have a hacker produce the receipts, has basically confirmed the thing it wouldn't say outright. Calling scraped YouTube rips and Genius lyrics "publicly available" is a legal argument, not a factual one, and courts tend to notice the difference. And quietly deciding a breach involving customer payment details doesn't meet the bar for notification is the kind of call that looks a lot worse in hindsight than it does in a press statement.

Read more about this at: The Verge

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.