The results show that VLM image tokens provide useful semantic cues for transferring driving attributes to objects outside the training vocabulary, and this model matches strong vision-only segmenters on familiar categories and improves transfer to real open-world anomalies.
Abstract
Driving in the real world is open-world: a car may encounter a fallen mattress, a deer, or other objects outside its training data. Naming them is not enough. The system must know how to treat each region: can it drive over it, and how severe would a collision be? We therefore shift scene perception from category labels to dense action-relevant attributes, where each pixel is labeled by how it should affect motion rather than by object name. We instantiate this general formulation with two ordered attributes: 7-rank drivability and 5-rank vulnerability. We read Qwen3.5 image-token hidden states directly as a spatial semantic representation. A lightweight boundary-aware decoder then turns this coarse token grid into sharp full-resolution attribute maps. The whole process requires neither autoregressive text generation nor an external mask model such as SAM. We train on dense attribute labels built in CARLA and test transfer to real scenes and to novel obstacles never seen in training. We compare with vision-only segmenters trained on the same attributes and prompted VLM segmenters. Our model matches strong vision-only segmenters on familiar categories and improves transfer to real open-world anomalies, reaching 69.4% mean vulnerability-rank recall versus 57.1% for the best vision-only baseline and 53.9% for the best prompted VLM baseline. These results show that VLM image tokens provide useful semantic cues for transferring driving attributes to objects outside the training vocabulary.
A decoupled semantic understanding framework that first resolves predefined traffic questions into structured semantic facts and subsequently uses these facts to guide caption generation, leading to more accurate and reliable traffic understanding, leading to higher-quality lan guage generation.
Bui Hoai Thuong Nguyen, Thanh-Nhan Vo, T. Nguyen et al.· 1 citation
Locating a specific object instance in a cluttered scene using a single reference image and a short description, and reporting when that instance is absent, large vision-language models usually address this task. We ask whether the same capability is available far more cheaply, from representations already learned by a...
K. Gupta, Ahmed Rafi Hasan, Md. Mahfuzur Rahman et al.· 0 citations
Extracting who is where, on which team, wearing which number from a broadcast frame is typically done by stitching a detector, an OCR engine, and classifiers together -- and the stitching step swaps identities under occlusion. We make association native instead: a 0.77B vision-language model (Florence-2) is fine-tuned...
A single consumer-grade GPU running a vision-language model (VLM) is used to supply missing guidance on pixel-level segmentation, improving segmentation while producing structured, auditable evidence that drives the result and can be inspected on its own.
Teresa DiMeola, Charles Walter, Hong Xiao· 0 citations
Long-term change understanding from images of the same place revisited over time is a challenging task with applications in map maintenance and urban infrastructure monitoring. Prior work addresses it either through pixel-level prediction or difference captioning, neither of which is sufficient to reliably measure how...
Benedetta Liberatori, Nermin Samet, Paolo Rota et al.· 0 citations
This work addresses open-world semantic segmentation, the joint task of segmenting known classes while detecting and grouping novel or anomalous content without additional supervision, by extending a dual-decoder baseline with a third, complementary decoder within a unified encoder-decoder design.
Anastasios Romanos Varvarigos, Nikos Giakoumoglou, Tania Stathaki· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.