AI
Meta details Muse Realtime Avatar, an audio-driven video model built for live calls
Meta’s AI research blog describes Muse Realtime Avatar as an audio-driven Diffusion Transformer conditioned on a speech-token stream, reference media and a rolling window of recent video latents. It generates short causal chunks so the newest output becomes motion context for the next, keeping voice, lip motion and expression synchronised for the length of a conversation.