vLLM-Omni introduces unified serving for Audio, Video, DiTs, and Vision
Key point
The new framework extends vLLM's PagedAttention to Diffusion Transformers and supports models like Qwen3-Omni and Wan2.2.
Details
vLLM-Omni addresses the fragmentation of multimodal AI serving by unifying text, audio, diffusion, and vision models into a single inference framework. Previously, serving these mixed workloads required juggling multiple frameworks across separate ports or containers.
Core Architecture
The system brings PagedAttention and KV cache optimizations from the standard vLLM engine to Diffusion Transformers (DiTs) and parallel generation models. It utilizes a disaggregated pipeline called OmniConnector to pass data between stages across GPUs, enabling real-time full-duplex audio, video, and image generation.
Supported Models
The framework currently supports a range of multimodal and robotics models, including:
- Qwen3-Omni
- MiniCPM-o 4.5
- MiniMax H3
- Wan2.2
- Qwen3-TTS
- Robotics action models like $\pi_0$.5
The project provides OpenAI-compatible APIs and is available via GitHub and a technical report.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.