Expanding on how Voice Engine works and our safety research
OpenAI Blog
OpenAI described how its Voice Engine text-to-speech model generates realistic synthetic voices from brief audio samples and text inputs. The system requires only a 15-second voice sample to create a matching synthetic voice. The company emphasized safety considerations including detection systems for identifying synthetic speech and working with policymakers on responsible deployment guidelines.
Why it matters
Exploring the technology behind our text-to-speech model.