AI Briefing
KO

LongCat Video Avatar 1.5

·2026.05.23 12:27

Key point

The open-source audio-driven human video generation framework LongCat-Video-Avatar 1.5 has been released.

Details

The Meituan team has released LongCat-Video-Avatar 1.5, an open-source framework for audio-driven human video generation. It is an upgraded version focused on commercial-grade stability and extreme empirical optimization.

Key Improvements

• Whisper-Large Audio Encoder: Replaces the previous Wav2Vec2 to achieve much smoother and more natural lip synchronization • Commercial Stability: Supports long video generation while maintaining precise lip-sync, full-body temporal stability, and strict identity consistency • Stylized Domain Generalization: Generalizes robustly to complex real-world conditions including animation, animals, multi-person interactions, and object manipulation • Efficient 8-Step Inference: DMD2-based step distillation accelerates inference to 8 NFE, balancing cost-efficient serving with excellent visual fidelity

Supported Tasks: Natively supports Audio-Text-to-Video (AT2V), Audio-Text-Image-to-Video (ATI2V), and Video Continuation, and is compatible with both single-stream and multi-stream audio input.

A human evaluation benchmark was introduced, comprising a total of 508 image-audio source pairs across 6 application scenarios (news broadcast, knowledge education, daily life, entertainment, singing, commercial promotion), 2 languages (Chinese/English), and 2 visual styles (realistic/animated). 770 crowdsourced evaluators rated human-likeness on a 1-5 scale, collecting a total of 13,240 judgments, and 10 domain experts conducted structured quality analysis across 4 dimensions.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.