TLDRocket
Sign in

ByteDance Seed Introduces SeedRealtime: a Native Audio-Visual Full-Duplex LLM That Watches, Listens and Speaks in One Model

MarkTechPost Michal Sutter

ByteDance showed SeedRealtime, one model that can watch, listen and talk at once. It’s already inside Doubao, but outsiders can’t plug into it yet.

Based on reporting by MarkTechPost, Michal Sutter — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

ByteDance’s Seed team has put out SeedRealtime, a single model that handles audio, video and text together instead of handing each step off to a different system. The company is pitching it as a move toward “omni-modal” interaction, and the key shift is simple: the model keeps up with a live stream, rather than waiting for neat turns the way most voice assistants still do.

That matters because the usual stack is a chain of separate parts — speech recognition, then vision, then text-to-speech — and every handoff can add delay or lose detail. SeedRealtime tries to keep perception, understanding, decision-making and expression inside one end-to-end system. Even turn-taking is handled inside the model, so it does not rely on an external voice-activity detector to decide when to answer.

Seed says the model shows three big gains: joint audio-visual understanding, proactive interaction and more natural conversational timing. The demos lean hard on that claim. In one, it links names to faces at a noisy dinner and keeps identities straight even when different people talk about different travel plans. In another, it is told to wait for a bronze screen stand at the Hebei Museum, then speaks up on its own when the object finally comes into view.

The other examples push the same point from different angles. It can follow a fast-flipping ResNet paper well enough to stop on the “3.4 Implementation” section and read out the learning rate, momentum and weight decay. It can also watch an espresso workflow, interrupt when whole beans are used by mistake, and suggest shortening extraction by 2 to 3 seconds.

There is one catch, and it is a big one for anyone outside ByteDance. SeedRealtime is live in the Doubao app, but ByteDance has not published a technical report, a parameter count, open weights or an endpoint through Volcano Engine or BytePlus. So the model exists, the idea is proven, and the rest of the ecosystem is still waiting at the door.

My take — AI-written commentary, not fact-checked reporting

This is the kind of closed-model flex that looks impressive and inconvenient at the same time. ByteDance gets to show off the future of real-time multimodal assistants while keeping everyone else outside the velvet rope. The industry keeps praising “openness” right up until a company proves the demo works, then suddenly everyone discovers the joys of secrecy and product lock-in.

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.