Aug 2026· 2026 IEEE International Conference on Mechatronics and Automation (ICMA)· pp. 1234-1239· 0 citations· 21 references
Abstract
To address the inherent limitations of Vision-Language Models in long-tail object retrieval for autonomous driving, this paper proposes a Dual-Granularity Structured Scene Retrieval (DG-SSR) architecture. By decoupling text queries and visual features, we introduce a parameter-free mechanism that fuses local semantic scores with macro global context. An Adaptive Negative Injection (ANI) strategy and a Soft-NCE loss further enforce fine-grained alignment and mitigate color bias. Evaluated on a curated nuScenes dataset comprising 8,500 homogeneous street views and 1,128 combinatorial queries, our method achieves an mP@5 of 33.06% with a single-query latency of 0.812 ms, outperforming CLIP (16.48%) and BLIP (23.88%) by significant margins. Extensive ablation analysis demonstrates that optimal retrieval in complex scenes is achieved through a local-dominated architecture supplemented by minimal global context.
Multimodal large language models (MLLMs) must balance local detail against scene context when interpreting ultra-high-resolution (UHR) remote sensing (RS) imagery within a limited visual-input budget. Existing selection-based methods either prune tokens and select patches through relevance scoring, or crop actively thr...
Yao Zhang, Peng-Yu Dai, Wei Guo et al.· 0 citations
Due to extreme density variations and scarce visual features, tiny-object detection in low-altitude uncrewed aerial vehicle (UAV) images remains a challenging task. Although DETR-based detectors have shown promising potential, they rely on a fixed number of object queries, which leads to suboptimal performance in dense...
Ling-Feng Lin, Bo-Xiang Xie, Wen Luo et al.· IEEE Geoscience and Remote S...· 0 citations
RS-Florence is proposed, a compact unified model that addresses remote sensing perception systems through a Prompt-Driven Sequence-to-Sequence framework, which maps images and task-specific prompts into a unified sequence of natural language and discrete geometric tokens.
Yang Liu, Wei-Xing Luo, Huai-Zhou Qi et al.· Proceedings of the Thirty-Fi...· 0 citations
Large collections of street-view imagery provide rich visual information about urban environments, but extracting fine-grained geographic information from such data remains challenging. In particular, fine-grained regional geolocalization is challenging because nearby areas often share coarse geographic cues. We study...
Chang-Le Lee, Yeonsoo Park, Abdullah Alfarrarjeh et al.· 0 citations
Current object detection models face significant challenges when applied to ultra-high-resolution images, as global downscaling frequently destroys crucial details for small objects, while EGC wastes computation on irrelevant regions and causes object truncation at tile boundaries. We propose a two-step reasoning archi...
Gia-Phuc Song-Dong, Anh-Kiet Tran-Nhu, Minh-Triet Tran et al.· International Conference on...· 0 citations
VAR has gained widespread popularity due to its next-scale prediction paradigm. However, it faces substantial performance bottlenecks when handling complex scenes with multiple objects and attributes. Existing diffusion-based enhancement methods fail to adequately address the unique challenge of cross-scale error propa...
Zhen-Nan Chen, Tianxing Shi, Pengcheng Xu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.