TLDRocket
Sign in

Jukebox

OpenAI

OpenAI built a neural net called Jukebox that writes full songs, including shaky singing, from scratch as raw audio. It even apes specific artists' styles, and they're giving away the code.

Based on reporting by OpenAI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

OpenAI dropped something genuinely strange this week: a neural network named Jukebox that writes music the hard way, sample by sample, rather than stitching together notes on a score. The result is raw audio in genres ranging from pop to hip-hop to country, complete with singing that sounds unmistakably human, if a little unsteady, like someone new to karaoke who nonetheless nails the melody.

What makes this different from earlier music-generating models is the level of the model actually works at. Most systems before it dealt in MIDI or sheet-music-style representations, leaving the messy job of turning notes into sound to a separate synthesizer. Jukebox skips that step entirely. It models audio waveforms directly, which is a much harder problem, since a few minutes of music can contain millions of individual samples. OpenAI's team got around that by compressing audio into a more manageable code using something called a VQ-VAE, then having a transformer generate music in that compressed space before decoding it back into sound.

The outputs aren't studio-ready. Anyone who listens will hear artifacts, mumbled lyrics, and moments where the model clearly loses the thread. But the fact it can mimic artist styles at all, producing something that sounds like it's reaching for a particular singer's phrasing or a genre's rhythmic feel, says a lot about how much structure these systems can absorb from raw sound alone, without anyone hand-coding rules about chord progressions or verse-chorus form.

OpenAI isn't keeping this behind a paywall or an API waitlist, at least for now. They're releasing the model weights and the code, plus a sample explorer so people can poke around at what the thing has already generated before running anything themselves. That openness is notable given how OpenAI's release strategy has shifted since, but at the time it fit a pattern of treating generative audio as a research curiosity worth sharing widely rather than a product to gatekeep.

My take — AI-written commentary, not fact-checked reporting

I'll take a slightly-broken open release over a polished closed demo every time, and Jukebox is a good reminder that OpenAI used to default to that instinct before GPT-3 changed the calculus. The singing is rough, sure, but rough and inspectable beats smooth and locked away, especially for something this experimental.

Read more about this at: OpenAI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.