TLDRocket
Sign in

Introducing GPT‑Live

Simon Willison's Weblog Simon Willison Covered by 6 sources

OpenAI just rolled out a new voice model for ChatGPT called GPT-Live, replacing the old GPT-4o-based voice mode. It can quietly hand off tough questions to GPT-5.5 mid-conversation without breaking the flow.

Based on reporting by Simon Willison's Weblog, Simon Willison — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

OpenAI has quietly replaced the brains behind ChatGPT's voice mode with something called GPT-Live, and if Simon Willison's few weeks of preview access are any indication, it's a meaningful jump. The old voice mode ran on a GPT-4o era model with a knowledge cutoff sometime in 2024, and Willison says he'd mostly given up on it because the model was too weak to be much of a brainstorming partner.

The interesting design choice here is what happens when a question is actually hard. Rather than trying to do everything itself in real time, GPT-Live can delegate web searches, deeper reasoning, or complex tasks to GPT-5.5 running behind the scenes, then fold the answer back into the conversation once it's ready. Crucially, the voice model keeps talking to you while that happens, so the conversation doesn't grind to a halt waiting for a slower model to think. OpenAI says it plans to keep swapping in newer frontier models behind GPT-Live as they ship, so the system is built to get smarter without a rename.

It wasn't all smooth during the preview, though. Willison ran into an odd bug where the model would interrupt him mid-sentence to laugh at things he hadn't said as jokes at all. He describes it as feeling rude and condescending, which is a strange complaint to have about a chatbot but a fair one. He flagged it to OpenAI, and after some apparent tweaks, the random laughing fits seem to have calmed down. Digging through his own transcripts, he traced one likely trigger to an offhand question about where owls hide during the day.

And despite the hiccup, Willison's real-world test is a decent sign of how far the new model has come: a full hour-long conversation while walking his dog, snapping photos of pelicans along the way. No owl photos yet, but the fact that he wanted to keep talking for an hour says more about GPT-Live's usefulness than any spec sheet would.

My take — AI-written commentary, not fact-checked reporting

The smart move here isn't the voice quality, it's the quiet handoff to a bigger model when things get hard, because that's the actual fix for why voice assistants have felt dumb for years. Everyone's been chasing snappier latency while ignoring that snappy answers to hard questions are usually wrong answers. OpenAI finally admitting that sometimes you just need to pause and think, without making the user feel the pause, is the kind of unglamorous engineering that actually moves the needle. The laughing bug is a reminder that these systems still don't really understand tone, they're just pattern-matching hard enough to fool us most of the time.

Read more about this at: Simon Willison's Weblog

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.