AI Briefing
KO

Meta Introduces Muse Realtime Avatar for Synchronized, Expressive Video Generation

·2026.09.23 09:00

Key point

Meta has introduced Muse Realtime Avatar, a system that generates synchronized, expressive video from speech tokens with subsecond latency, outperforming Runway and HeyGen in user preference tests.

1 / 12

Details

Meta has launched Muse Realtime Avatar, a technology that transforms Muse Realtime Voice into interactive, embodied avatars. The system operates as a single streaming pipeline where Muse Realtime Voice produces speech tokens (VQs) that drive both audio and video generation, ensuring tight synchronization between voice, lip motion, and facial expressions.

To achieve real-time performance, Meta employed a distillation technique that reduces the neural function evaluations per chunk from 120 (in the teacher model) to just 2 (in the student model), a 60x reduction in computational cost while maintaining high visual quality. The system streams 448x768 portrait video at 25 frames per second with approximately 870 ms of latency. Through optimizations like persistent KV caches, four-bit quantization-aware training, and NVIDIA CUDA Graph capture, Meta achieved an 8x increase in serving capacity, allowing 12 concurrent sessions on a single GB200 GPU.

In live conversation tests against leading commercial systems, raters preferred Muse Realtime Avatar over Runway Characters by 78% to 22% and over HeyGen LiveAvatar by 88% to 12%. The system includes safety measures such as Meta Video Seal, an invisible watermark embedded in all generated video to ensure traceability.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.