Shengshu unveils Vidu S1, a real-time voice-controlled video model
Key point
Vidu S1, an interactive video model that responds in real time to users' voice instructions and generates video indefinitely, has been unveiled.
Details
Vidu S1, jointly announced by Tsinghua University and Shengshu Technology, is an interactive video generation model that departs from the existing offline batch generation approach, allowing users to intervene with voice during generation to change the direction of the video in real time.
This model solves the drift (screen collapse) caused by error accumulation, a limitation of existing autoregressive models, and is designed so that video can continue indefinitely without interruption. It is also based on TurboDiffusion and TurboServe technology, achieving fast inference speed of 540p resolution and up to 42 FPS even on consumer GPUs.
Key features:
- Real-time interaction: Character actions can be controlled via voice commands at any point during generation
- Infinite-length generation: Supports stable streaming generation through error-accumulation mitigation technology
- High efficiency: Provides a real-time-level framework even in low-cost GPU environments
- Support for diverse characters: Custom characters can be applied, from real people to animations and pets
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.