Muse Realtime Avatar brings expressive characters to Meta’s Muse
Meta AI Research ● Covered by 14 sources
Meta’s Muse can now turn a voice chat into a live avatar. It’s aiming for real-time, expressive video without the usual lag or drift.
Based on reporting by Meta AI Research — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Meta is adding a new layer to Muse: not just speech, but a moving face, body, or even an object that reacts in real time. The system is called Muse Realtime Avatar, and it takes the same speech token stream used by Muse Realtime Voice to keep words, lip motion, and expression in sync.
The pitch is broad. A portrait can show subtle expressions, an illustrated character can gesture and shift posture, and animals or everyday objects can still look like themselves while behaving as if they’re in a conversation. Meta says the avatar stays coherent from one turn to the next, instead of resetting its look and mannerisms every few seconds like so many demo systems do.
Under the hood, Muse Realtime Avatar is an audio-driven Diffusion Transformer conditioned on speech tokens, reference media, and a rolling window of recent video latents. It generates video in short causal chunks, then feeds the newest latents forward as context for the next chunk. That is the trick that keeps the system moving without letting quality fall apart as the conversation goes on.
Meta also says it cut the model from a 120-evaluation teacher setup down to an unguided two-step student, a 60x reduction, while keeping quality close to the original. In live tests against Runway Characters and HeyGen LiveAvatar, raters preferred Muse Realtime Avatar overall and on every evaluated dimension, though the mannerism comparison with Runway was not statistically different from parity.
The serving side matters too. Meta says the system runs at 448x768 portrait video, 25 frames per second, with about 870 ms of latency from the end of a user turn to the first byte of the synchronized response. On a single GB200, that translates to 12 concurrent real-time sessions for video generation, helped by cache-aware routing, dynamic batching, quantization-aware training, and work with NVIDIA on model optimizations. Meta is also embedding an invisible watermark with Video Seal, which is the least flashy part of the whole thing and probably the most necessary one.
My take — AI-written commentary, not fact-checked reporting
This is the right direction: embodied AI only becomes useful when the face, voice, and timing stop fighting each other. The impressive part isn’t that Meta can make a cartoon talk; it’s that it’s trying to make the whole stack behave like one system instead of three demos in a trench coat. The watermarking bit deserves more attention too, because once every object can speak, traceability stops being a nice extra and starts being table stakes.
Read more about this at: Meta AI Research
Related stories
Meta Introduces Muse, a Personal AI Agent That Runs on Its Own Dedicated Secure Cloud Computer
MarkTechPost · 2 weeks ago ·
42
Meta Superintelligence Labs Releases Muse Voice Transcribe: One Real-Time Model for Streaming ASR, Diarization, and Endpointing
MarkTechPost · 3 weeks ago ·
5