This work proposes a segmentation framework guided by Contrastive Language–Image Pre-training (CLIP) that enriches sparse 3D tokens with vision–language semantic priors and introduces a decoupled CLIP-induced semantic residual that forms semantic-geometric attention biases for local window attention.
Abstract
Recent 3D Transformers have become a dominant framework for point-cloud segmentation by modeling spatial context in sparse 3D scenes. However, geometry and color alone provide limited high-level semantic cues, especially for cluttered boundary regions, visually similar objects, and long-tail categories. To address this issue, we propose a segmentation framework guided by Contrastive Language–Image Pre-training (CLIP) that enriches sparse 3D tokens with vision–language semantic priors. Specifically, dense CLIP features extracted from multi-view RGB images are projected onto 3D points through visibility-aware alignment and view pooling, and are fused with relative geometric offsets and color cues to form semantically aware sparse voxel tokens. To better exploit the aligned CLIP semantics during local token interactions, we build on contextual relative signal encoding (cRSE) and introduce a decoupled CLIP-induced semantic residual that forms semantic-geometric attention biases for local window attention. We further adapt block-wise online softmax computation to generate and consume these biases on the fly. Experiments on ScanNet, ScanNet200, and S3DIS demonstrate competitive segmentation performance, improved instance-level discrimination, and a 25.7% reduction in peak
online 3D-stage
training memory compared with the materialized attention implementation when cached CLIP features are used.
3D Gaussian Splatting provides an efficient representation for 3D reconstruction, and recent extensions attach semantic attributes to Gaussians for open-vocabulary scene understanding. However, lifting view-dependent 2D foundation-model outputs into 3D space introduces cross-view inconsistencies and weak geometric grou...
Vision-language models excel at 2D image understanding but remain limited in 3D spatial reasoning. Progress is hindered by limitations in current benchmarks. First, 3D datasets often rely on point clouds that capture geometry but discard rich visual features like texture, text, and materials. Second, annotations treat...
Anubhav Khanal, Prabigya Acharya, Roshni Poudel et al.· 0 citations
Unified 3D vision-language systems must combine complementary geometry, scale, and appearance cues while supporting tasks from instance segmentation to language-guided reasoning. Existing methods often process point clouds, voxel grids, and multi-view images independently; directly combining these heterogeneous represe...
Xue-Qi Qiu, Xing-Yu Miao, Jing-Jing Deng et al.· 0 citations
A 3D Local- Global Linear Attention Mechanism (LG-LAM) is devised that efficiently captures long-range contextual information with linear complexity, enabling a comprehensive understanding of the 3D scene without heavy computational burdens.
Jie Li, Jia-Heng Xu, Laiyan Ding et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.