TLDRocket
Sign in

Expanding on how Voice Engine works and our safety research

OpenAI

OpenAI detailed Voice Engine, its tool that clones a voice from just 15 seconds of audio. It's staying in a small testing group for now — the company says the deepfake risk is too real to ship wide.

Based on reporting by OpenAI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

OpenAI pulled back the curtain a bit on Voice Engine, the text-to-speech system it's been sitting on since late 2022. Feed it a 15-second clip of someone talking and a script, and it spits out speech that sounds like that person, complete with their accent and emotional inflection. That's a tiny sample size for something this convincing, and it's exactly why OpenAI is treating the whole thing like a live wire.

Instead of a public launch, the company has kept Voice Engine in the hands of around 10 partners since last year, testing use cases like reading assistance for non-readers, translation that preserves a speaker's original voice, and support tools for people who've lost their ability to speak. One partner, Age of Learning, is using it for education content. Another effort involves recreating speech for a nonprofit that helps nonverbal users communicate. These are the sympathetic cases OpenAI wants people to picture first.

But the company isn't pretending the obvious problem doesn't exist. A tool that can mimic anyone's voice from a few seconds of audio is a gift to scammers, political dirty-tricksters, and anyone running an impersonation scheme — and 2024, an election year in dozens of countries, is a rough time to hand that out freely. OpenAI says it requires explicit consent from the original speaker before cloning a voice, discloses when audio is AI-generated, and has built watermarking into the output so synthetic speech can be traced back. Partners also have to agree not to use the tool on voices without permission.

OpenAI's framing here is deliberately cautious, almost defensive — this is a company that got burned by comparisons to voice-cloning scams and wants to show its homework before wider release. Whether that satisfies regulators or just delays the inevitable is the open question. The underlying tech clearly works well enough to be dangerous, which is precisely why they're not rushing to hand out the keys.

My take — AI-written commentary, not fact-checked reporting

Watermarking and consent clauses are the right instincts, but they're paper barriers against a technology this good — once a comparable open model shows up, and it will, none of those safeguards travel with it. OpenAI slow-walking release looks responsible today, but it mostly buys time rather than solving the problem, and I'd rather see it paired with real detection infrastructure than another blog post about good intentions.

Read more about this at: OpenAI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.