inclusionAI Releases Real-Time Multimodal Dialogue Model
Key point
inclusionAI has released Realtime-Venus, a 9B-scale multimodal dialogue model capable of processing video and audio in real time.
Details
inclusionAI has released two checkpoints of the real-time multimodal dialogue system Realtime-Venus via Hugging Face. Developed based on MiniCPM-o 4.5, this model supports proactive interaction, allowing it to perceive situations and respond first without requiring user prompts.
Key Features
- Native full-duplex dialogue: Continuously perceives the surroundings while speaking, distinguishing and naturally responding to user backchannels, interruptions, and corrections.
- Proactive interaction: Continuously processes temporally aligned video and audio to initiate dialogue first when an event requiring a response occurs, without waiting for user input.
- Training-free long-video memory: Understands long video content without additional training by storing important visual information and retrieving non-duplicate materials relevant to queries to reconstruct context.
- Asynchronous Delegation: Requests external backend tasks within the stream and processes results asynchronously to avoid disrupting the dialogue flow.
Model Composition
The released models are available in two versions.
- Realtime-Venus-Omni: A 9B-scale audio-visual interaction model that simultaneously generates text and voice output.
- Realtime-Venus-Audio: A model specialized for audio processing, supporting audio-based dialogue and text/voice output.
Both models are available for download on Hugging Face, and Realtime-Venus-Harness, a runtime for external tool integration, is provided in the GitHub repository.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.