Qwen-Drive-1.0 Launches as a Unified AI Model Combining 3D Perception and Motion Planning for Autonomous Driving
Summary
Qwen-Drive-1.0 launches as a powerful unified AI model combining 3D perception, visual question answering, and motion planning for autonomous driving, outperforming competitors on major benchmarks and now freely available under an open-source license.
Key Points
- Qwen-Drive-1.0 is a new vision-language model for autonomous driving that builds on the Qwen3.5-4B architecture, integrating 3D perception, visual question answering, and motion planning into a unified framework.
- The model features two external modules — a BEV Perception Head for 3D object detection and scene segmentation, and a Planning Expert for generating future driving trajectories — while preserving the base model's general vision-language capabilities.
- Qwen-Drive-1.0-SFT outperforms competing models across driving VQA benchmarks, achieving top scores on LingoQA, SURDS, and WaymoQA, and the model is publicly available on Hugging Face and ModelScope under the Apache 2.0 license.