TLDRocket
Sign in

Gradium Launches Voice Design: Write a Prompt, Get a Brand New Synthetic Voice in Seconds

MarkTechPost Michal Sutter

Gradium now makes synthetic voices from a text prompt in seconds. That skips voice cloning, and it’s free on every plan, even the free tier.

Based on reporting by MarkTechPost, Michal Sutter — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Gradium thinks the problem with voice work is not lack of models. It’s lack of the right voice. A catalog can have 400 options and still miss the one a team actually needs: the Quebecois receptionist, the older narrator with lecture-hall weight, the accent that sounds local instead of generic. The Paris-based company, spun out of Kyutai, is answering that gap with Voice Design, a tool that turns a written description into a brand new synthetic voice in a few seconds.

The pitch is simple and a little provocative: no reference audio, no speaker to clone, no rights clearance dance. The only input is the description. Gradium says the system responds to the usual casting-language traits — gender, age band, accent or origin, pitch, pace, energy, timbre, resonance, register, manner, and the job the voice is meant to do. Prompts can run from 1 to 500 characters and are accepted in English, French, Spanish, Portuguese, or German. The company also advises ending with the intended use, because that steers delivery and register, not just the sound.

Voice Design is already live in the Gradium API and in Studio, and it’s free on every plan, including the free tier. A request returns 1 to 5 candidates, usually in 3 to 5 seconds. From there, the workflow is ordinary enough: poll until the candidate is ready, audition it through the standard text-to-speech endpoint, then promote the one you want. Once converted, the voice behaves like any other kept voice in the system and uses the same streaming endpoint, latency, and output formats.

There are catches, but they’re practical rather than dramatic. Unconverted candidates vanish after 30 days. They can only be auditioned through REST, with text capped at 100 characters, and they do not work over the TTS WebSocket or Speech-to-Speech. Converting a candidate is free and uses one custom voice slot, shared with cloned voices. The free tier holds 5 custom voices; paid plans hold 1,000.

Gradium is also making a point about quality. In a blind pairwise listening test across six voice design systems and five languages, with 7,627 comparisons, it says Voice Design won 72.6% of the time. That beat ElevenLabs at 59.0%, then Inworld, Fish Audio, and MiniMax. The biggest gaps were on regional accents that catalogs usually flatten: Quebecois French, Rioplatense Spanish, Bavarian German, Colombian Spanish, and African Portuguese. A separate model judge, Gemini 3.1 Pro, gave the same overall ranking. Vendor-run benchmarks are never a love letter to humility, but this one is at least trying to measure the thing buyers actually care about: does the voice sound right, or just plausible?

My take — AI-written commentary, not fact-checked reporting

This is the kind of product that makes cloning look like yesterday’s awkward workaround. Gradium is betting that teams want promptable voices, not a rights-management hobby. In Europe especially, that’s a cleaner story than pretending every synthetic voice needs to start with someone else’s throat.

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.