Skip to content
Open access

Vision–language guided semantic-geometric transformer for memory-efficient 3D scene understanding

Sep 2026 · Frontiers in Artificial Intelligence · Vol 9 · 0 citations · 55 references

TL;DR

This work proposes a segmentation framework guided by Contrastive Language–Image Pre-training (CLIP) that enriches sparse 3D tokens with vision–language semantic priors and introduces a decoupled CLIP-induced semantic residual that forms semantic-geometric attention biases for local window attention.

Abstract

Recent 3D Transformers have become a dominant framework for point-cloud segmentation by modeling spatial context in sparse 3D scenes. However, geometry and color alone provide limited high-level semantic cues, especially for cluttered boundary regions, visually similar objects, and long-tail categories. To address this issue, we propose a segmentation framework guided by Contrastive Language–Image Pre-training (CLIP) that enriches sparse 3D tokens with vision–language semantic priors. Specifically, dense CLIP features extracted from multi-view RGB images are projected onto 3D points through visibility-aware alignment and view pooling, and are fused with relative geometric offsets and color cues to form semantically aware sparse voxel tokens. To better exploit the aligned CLIP semantics during local token interactions, we build on contextual relative signal encoding (cRSE) and introduce a decoupled CLIP-induced semantic residual that forms semantic-geometric attention biases for local window attention. We further adapt block-wise online softmax computation to generate and consume these biases on the fly. Experiments on ScanNet, ScanNet200, and S3DIS demonstrate competitive segmentation performance, improved instance-level discrimination, and a 25.7% reduction in peak online 3D-stage training memory compared with the materialized attention implementation when cached CLIP features are used.

Read PDF

Similar papers

Preprint Aug 2026

GaussianDS: Depth-supervised Semantic Gaussian Splatting for Scene Understanding

3D Gaussian Splatting provides an efficient representation for 3D reconstruction, and recent extensions attach semantic attributes to Gaussians for open-vocabulary scene understanding. However, lifting view-dependent 2D foundation-model outputs into 3D space introduces cross-view inconsistencies and weak geometric grou...

Yu-Fei Zhang, Chen-Lu Zhan, Hong-Wei Wang · 0 citations
#artificial intelligence Preprint Sep 2026

SceneBench: A Hierarchical Benchmark for Vision-Language Understanding of 3D Scenes

Vision-language models excel at 2D image understanding but remain limited in 3D spatial reasoning. Progress is hindered by limitations in current benchmarks. First, 3D datasets often rely on point clouds that capture geometry but discard rich visual features like texture, text, and materials. Second, annotations treat...

Anubhav Khanal, Prabigya Acharya, Roshni Poudel et al. · 0 citations
Preprint Sep 2026

From Alignment to Fusion in 3D Vision-Language

Unified 3D vision-language systems must combine complementary geometry, scale, and appearance cues while supporting tasks from instance segmentation to language-guided reasoning. Existing methods often process point clouds, voxel grids, and multi-view images independently; directly combining these heterogeneous represe...

Xue-Qi Qiu, Xing-Yu Miao, Jing-Jing Deng et al. · 0 citations
Conference Aug 2026

Enhancing 3D semantic scene completion via efficient attention and feature augmentation

A 3D Local- Global Linear Attention Mechanism (LG-LAM) is devised that efficiently captures long-range contextual information with linear complexity, enabling a comprehensive understanding of the 3D scene without heavy computational burdens.

Jie Li, Jia-Heng Xu, Laiyan Ding et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.