SeedRealtime: ByteDance Unveils Full-Duplex AI
Key point
ByteDance Seed has deployed SeedRealtime, a full-duplex model that processes audio and video in real time, on Doubao.
Details
ByteDance Seed has unveiled SeedRealtime, a full-duplex model that simultaneously understands audio and video, listens and watches while the user is speaking, and determines when to speak on its own.
Existing voice AI typically uses a half-duplex structure where VAD detects the end of the user's speech before responding. SeedRealtime integrates audio/video input, conversation timing judgment, and voice output into a single end-to-end model, targeting the following capabilities:
- Audio-visual combined understanding leveraging screen information
- Proactive interaction based on scene changes or context
- Natural conversation timing that distinguishes background noise and chatter, allowing for interruptions or waiting
ByteDance Seed stated that in its own human evaluation, conversation timing issues such as speech interruptions, response delays, and misinterpretations of background noise were reduced by half compared to cascade approaches. However, weights, technical reports, and public APIs were not provided, and the model was released as a service deployed on ByteDance's Doubao app.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.