Preprint
Aug 2026
VOLA: Improving Open-World Driving by VLM-Based Semantic Attribute Prediction
The results show that VLM image tokens provide useful semantic cues for transferring driving attributes to objects outside the training vocabulary, and this model matches strong vision-only segmenters on familiar categories and improves transfer to real open-world anomalies.
Yu-Chen Zhang, Yuan Gao, Sebastian Schmidt et al.
· 0 citations