Qwen-Drive-1.0, Vision-Language Foundation Model for Autonomous Driving Released
Key point
The Qwen team released Qwen-Drive-1.0, a vision-language foundation model dedicated to autonomous driving that integrates 3D perception and motion planning.
Details
The Qwen team released Qwen-Drive-1.0, the first vision-language foundation model designed for autonomous driving. This model integrates 3D perception, visual question answering (VQA), and motion planning by attaching external modules without modifying the pre-trained VLM architecture.
Core Architecture and Features
- Base Model: Based on the native multimodal Qwen3.5-4B, maintaining the pre-trained VLM structure as is.
- BEV Perception Head: Acts as an explicit and inspectable 3D probe that processes multi-view inputs to perform 3D object detection, semantic occupancy prediction, and BEV map segmentation.
- Planning Expert: A diffusion transformer aligned with VLM representations, generating 5-second ego trajectories via flow matching.
Performance and Results
Qwen-Drive-1.0-SFT achieved an average score of 69.43 on autonomous driving VQA, surpassing general-purpose VLMs and autonomous driving-specific models. It also demonstrated competitive performance on major benchmarks such as LingoQA and WaymoQA, effectively adapting to the autonomous driving domain while maintaining general vision-language capabilities (MMBench, MMMU, etc.).
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.