TLDRocket
Sign in

Introducing next-generation audio models in the API

OpenAI

OpenAI just added new audio models to its API. Now you can tell text-to-speech how to sound, not just what to say.

Based on reporting by OpenAI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

OpenAI has pushed a fresh set of audio models into its API, and the headline feature is a small one that changes a lot: you can now tell the text-to-speech engine how to perform, not just what to read aloud. Type something like "talk like a sympathetic customer service agent" and the model adjusts tone, pacing, and warmth accordingly, rather than defaulting to the same flat narrator voice every time.

That sounds like a minor tweak, but it's the difference between a voice bot that reads a script and one that sounds like it's actually listening. Anyone who has sat through an automated phone tree knows the gap between synthetic speech and something that feels remotely human. Steerable delivery closes part of that gap without requiring a developer to record dozens of voice actors or fine-tune a separate model for every use case.

For companies building voice agents, this is the kind of control that used to require expensive custom voice work or clunky SSML tagging. Now it's a plain-language instruction, which means a support team can prototype a dozen personas in an afternoon and pick whichever one tests best with real customers. A patient tutor, a brisk dispatcher, a cheerful concierge — all from the same underlying model, steered by a sentence of prompt text.

The bigger story here is where OpenAI is placing its bets. Text and image generation get most of the attention, but voice interfaces are quietly becoming the default way people interact with AI outside a chat window — car assistants, phone support, smart speakers. Making the voice layer more expressive and controllable is less flashy than a new flagship chatbot, but it's arguably more useful for the businesses actually shipping products right now.

My take — AI-written commentary, not fact-checked reporting

Steerable voice is the unglamorous upgrade that actually matters, because most people don't talk to AI through a text box, they talk to it through a speaker or a phone call. I'd rather see this kind of practical control than another leaderboard-topping benchmark score, and I'll believe the hype about voice agents once support lines stop sounding like they're reading from a script written in 2009.

Read more about this at: OpenAI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.