Skip to content
Preprint

GaussianDS: Depth-supervised Semantic Gaussian Splatting for Scene Understanding

Aug 2026 · 0 citations · 40 references
Computer Science

TL;DR

GaussianDS, a depth-supervised semantic 3DGS framework that treats semantic lifting as a supervision-alignment problem and jointly optimizes RGB appearance, rendered depth, and compact semantics from scratch, is proposed.

Abstract

3D Gaussian Splatting provides an efficient representation for 3D reconstruction, and recent extensions attach semantic attributes to Gaussians for open-vocabulary scene understanding. However, lifting view-dependent 2D foundation-model outputs into 3D space introduces cross-view inconsistencies and weak geometric grounding, leading to severe semantic drift and boundary leakage. We propose GaussianDS, a depth-supervised semantic 3DGS framework that treats semantic lifting as a supervision-alignment problem and jointly optimizes RGB appearance, rendered depth, and compact semantics from scratch. Specifically, GaussianDS organizes unordered multi-view images into a pose-aware pseudo-video trajectory to propagate view-consistent masks via SAM2. During joint optimization, scale-shift-aligned monocular depth supervision and depth total-variation regularization stabilize Gaussian geometry, while a depth-edge-aware refinement loss explicitly anchors semantic transitions onto physical geometric discontinuities. Extensive evaluations show that our end-to-end framework not only retains high-fidelity 3D reconstruction and real-time rendering, but also establishes superior semantic understanding. GaussianDS sets new state-of-the-art performance on LERF (60.5% mIoU) and 3D-OVS (97.79% mIoU, 90.28% mBIoU) by mitigating semantic leakage, while seamlessly facilitating downstream 3D object removal.

View source

Similar papers

Open access Sep 2026

Semantic-Guided Adaptive Gaussian Segmentation

3D Gaussian Splatting (3DGS) enables real-time photorealistic scene reconstruction, yet its segmentation tasks suffer from two critical flaws: poor 3D consistency (e.g., blurred instance boundaries and unstable cross-view semantic association) and insufficient structural awareness near ambiguous object boundaries. To a...

Ying-Han Zhou, Fan Zhou · 0 citations
Preprint Sep 2026

D3GS: Depth, DINO, and RGB Diffusion Co-Guided 3D Gaussian Splatting for Sparse-View Reconstruction

Novel view synthesis from sparse inputs remains challenging for 3D Gaussian Splatting (3DGS) due to ambiguous geometry, cross-view inconsistency, and missing details in under-constrained regions, resulting in degraded reconstruction and unstable rendering. To tackle these issues, we propose D$^{3}$GS, a Depth-DINO-Diff...

Yun-Qi Gao, Zhan-Feng Liao, Han-Zhang Tu et al. · 0 citations
Open access Aug 2026

Semantic-guided 3D Gaussian splatting for sparse-view reconstruction in industrial digital twins

A semantic-guided 3D Gaussian splatting (3DGS) framework tailored to sparse-view industrial reconstruction was introduced, enabling robust reconstruction from limited viewpoints and offers a practical geometric foundation for automated inspection and remote equipment monitoring.

Boyang Li, Tian-Han Gao, Zuan Gu et al. · 0 citations
Preprint Sep 2026

NormLift: From Lifted Features To Semantic Reliability In 3D Gaussian Splatting

Training-free weighted aggregation is widely used to lift 2D semantic features onto 3D Gaussians for open-vocabulary scene understanding, yet its theoretical role remains insufficiently understood. Existing analyses typically justify this operation from the rendering side, treating Gaussian features as linearly composa...

Yi-Han Zang, Da Li, D. Engel et al. · 0 citations
Preprint Aug 2026

OutLangSplat: 3D Language Gaussian Splatting for UAV Outdoor Scenes

OutLangSplat is presented which adapts language Gaussian representations to UAV outdoor scenes by improving feature representation and aggregation reliability, and is the first accessible dataset of open-vocabulary 3D scene understanding for UAV outdoor scenes.

Xiaosheng Yan, Hefeng Wu, Yanghui Xu et al. · 1 citation
Preprint Aug 2026

FlexSplat: Flexible Feed-Forward 3D Gaussian Splatting without Point Cloud Correspondence

FlexSplat matches or approaches posed state-of-the-art reconstructors while requiring neither camera poses nor ground-truth depth, and matches the best perceptual (LPIPS) quality among the compared methods on GSO.

Amir Sabbaghziarani, Han-Ting Ye, Maria Gorlatova et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.