TLDRocket
Sign in

These execs think voice AI hasn’t reached its ChatGPT moment yet

TechCrunch Ivan Mehta

Voice AI is getting better, but these execs say it still hasn’t had its ChatGPT moment. The hard part now is speed, trust, and sounding human enough to keep people on the line.

Based on reporting by TechCrunch, Ivan Mehta — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

The voice AI pitch is loud right now. Investors have poured billions into startups across the space, from model makers to customer-service bots, meeting note-takers, and dictation tools. Every week brings a new release that claims it can talk like a person. But the people building it are not ready to declare victory.

PolyAI CTO Shawn Wen says voice AI still hasn’t hit its “ChatGPT moment,” even with full-duplex models now able to speak while listening. He said at the HumanX conference last month that the next hurdle is making reasoning fast enough that answers arrive quickly and the conversation feels natural. For customer service, that means an agent that doesn’t sound robotic and gives callers enough confidence to stay with it.

Wen’s view is that once the voice quality is good enough and a customer gets through the first few turns, trust starts to build. Then the bar changes. The goal is no longer just a polished demo; it is getting people to think they may not need a human if the AI can actually solve the problem.

Otter CMO Alex Gay put the emphasis somewhere else: speaker identification, intent capture, and combining that with organizational knowledge. Otter is also working on digital twins for meetings, and Gay said those avatars need to carry the same emotional cues as a real conversation. Without that, he said, it’s just a question-and-answer chatbot in a different outfit.

The weak spot keeps coming back to transcription. Wen said ASR systems often miss important keywords, and Gay agreed that Otter keeps pushing on accuracy because bad transcription poisons everything downstream. If the transcript is wrong, the follow-up actions are wrong too. And once the software starts taking action on the basis of a bad record, trust goes out the window. Both companies also stressed transparency: people should know when they’re talking to AI, and Otter even wants to notify people in chat when a meeting is being recorded.

My take — AI-written commentary, not fact-checked reporting

Voice AI is still in the embarrassing teenage phase: impressive enough to demo, not reliable enough to trust. The obsession with sounding human is almost beside the point if the transcript misses the key word and the follow-up goes off a cliff. Enterprises don’t need a charming bot; they need one that doesn’t lie by accident.

Read more about this at: TechCrunch

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.