OutLangSplat is presented which adapts language Gaussian representations to UAV outdoor scenes by improving feature representation and aggregation reliability, and is the first accessible dataset of open-vocabulary 3D scene understanding for UAV outdoor scenes.
Abstract
3D Language Gaussian Splatting embeds open-vocabulary language features into 3D Gaussian Splatting, providing an efficient explicit representation for text-driven 3D scene understanding. However, existing methods are limited to indoor or small-scale scenes, and tend to fail in Unmanned Aerial Vehicle (UAV) outdoor scenes, where severe occlusions and long distance viewpoints often lead to incorrect semantic activations and missing target responses. In this paper, we present OutLangSplat which adapts language Gaussian representations to UAV outdoor scenes by improving feature representation and aggregation reliability. For the feature representation, a 2D-3D dual-branch representation with region-based alignment and fusion is designed to improve spatial consistency, reducing incomplete target responses and background misactivations. For the feature aggregation, we introduce a training-free contribution and consistency-aware Gaussian feature aggregation strategy that leverages pixel contribution reliability and cross-view semantic consistency to suppress unreliable responses from noisy viewpoints. A new dataset is provided by manually annotating various objects on four real-world public UAV outdoor scene datasets. To the best of our knowledge, it is the first accessible dataset of open-vocabulary 3D scene understanding for UAV outdoor scenes. Quantitative evaluations and ablation studies demonstrate that OutLangSplat outperforms SOTA methods on both open-vocabulary semantic segmentation and instance localization tasks. The datasets and codes will be open-sourced.
In outdoor scene reconstruction, dynamic occlusions and multiscale structures often undermine multiview consistency and hinder effective gradient accumulation of high-frequency Gaussian primitives, leading to artifacts and the loss of fine details in Gaussian splatting–based radiance field methods. To address these challenges, we propose an uncertainty-aware hierarchical Gaussian splatting framework for outdoor 3D reconstruction. Specifically, our method constructs a hierarchical octree-based spatial representation from the results of aerial triangulation. It introduces level of detail constraints to enable structured management and progressive optimization of Gaussian primitives across different scales. This design effectively alleviates the imbalance in training and the redundant growth of Gaussian primitives commonly observed in multiscale outdoor scenes. In addition, we incorporate an uncertainty prediction mechanism that evaluates the consistency between rendered results and ground-truth images in the feature space, allowing the model to automatically identify dynamically occluded regions and suppress their gradient contributions during optimization. As a result, the adverse impact of dynamic artifacts on static scene modeling is substantially reduced. Experimental results demonstrate that, without incurring significant additional training overhead, our method consistently improves structural consistency and fine-detail reconstruction quality in outdoor scenes, while simultaneously reducing model complexity and maintaining real-time rendering performance. Furthermore, the proposed approach can be seamlessly integrated into multiple mainstream Gaussian splatting frameworks, exhibiting strong robustness and promising potential for practical deployment.
Junxing Yang, Haoran Gao, Chunyu Huang et al.· Journal of Electronic Imagin...· 0 citations
3D Gaussian splatting (3DGS) provides an efficient and expressive scene representation by jointly modeling spatial geometry and appearance, which has led to significant advances in high-fidelity scene reconstruction and novel view synthesis. However, in real-world environments, occlusions from pedestrians or equipment often introduce blurriness, artifacts, and geometric distortions. To address these challenges, this paper proposes a robust 3DGS modeling method enhanced by semantic and geometric consistency. First, the self-supervised foundation model DINOv2 is utilized to extract high-dimensional semantic features, leveraging its superior generalization capabilities to assist in identifying potential dynamic regions. Second, monocular depth estimation and a depth residual mechanism are introduced to construct geometric consistency constraints, enabling the precise localization of areas that violate static assumptions. Finally, a progressive guided probability masking mechanism is designed; it employs an adaptive sigmoid function to achieve a ‘coarse-to-fine’ soft-constraint optimization, effectively mitigating the training instability inherent in traditional binary hard masks. Experimental results on the neural radiance fields (NeRF)-on-the-go, RobustNeRF, and self-collected datasets demonstrate that the proposed method effectively suppresses dynamic artifacts and improves reconstruction quality. The proposed approach achieves competitive or superior performance compared with 3DGS, SpotLessSplats, T-3DGS, and RobustSplat on standard image-quality metrics, including peak signal-to-noise ratio, structural similarity index measure, and learned perceptual image patch similarity.
Wen Zheng, Guo Bao, Wenda Wang et al.· Engineering Research Express· 0 citations
Camera-reconstructed 3D scenes often require offline visual cleanup before inspection, presentation, or reuse as renderable virtual-scene assets. Representative applications include removing temporary furniture, parked vehicles, equipment, signage, and other distracting or obsolete objects from reconstructed indoor and outdoor environments. Such editing requires not only accurate target localization across viewpoints but also plausible recovery of the previously occluded background. Existing methods often depend on manually specified masks or category-restricted detectors, while projection-based pipelines independently inpaint multiple views and subsequently refine the 3D representation, potentially introducing cross-view appearance and geometry inconsistencies. We present OVR-GS (Open-Vocabulary Removal in Gaussian Splatting), an instruction-driven object-removal framework for pre-trained 3D Gaussian Splatting (3DGS) scenes. Given a free-form instruction, a language parser generates target-oriented queries and a textual background-completion condition. Grounding DINO and the Segment Anything Model (SAM) produce multi-view candidate masks, which are filtered using Contrastive Language–Image Pre-training (CLIP). The proposed Semantic-Aware Gaussian Selector (SAGS) aggregates rendering-contribution-weighted mask evidence, groups spatially coherent candidates, and identifies the target Gaussian subset through rendered-cluster semantic verification. After removal, new Gaussians are initialized from boundary-adjacent primitives and interior samples and optimized locally using Score Distillation Sampling (SDS), while the original background remains fixed. On IMFine, SPIn-NeRF, and Inpaint360GS, OVR-GS achieves peak signal-to-noise ratio (PSNR) values of 19.78, 17.82, and 24.62 dB and Fréchet inception distance (FID) values of 142.30, 148.60, and 34.80, respectively. The results demonstrate the effectiveness of localized Gaussian optimization for instruction-driven cleanup of reconstructed environments before visual inspection, presentation, or reuse as renderable virtual-scene assets.
Yongpeng Ding, Feng Ouyang, Jiawei Fan et al.· Italian National Conference...· 0 citations
ZeroSplat lifts 2D Vision-Language Model priors into 3D space through robust multi-view geometric constraints and enables intrinsic point-level understanding without incurring any additional feature storage, and significantly outperforms state-of-the-art methods across generalized and single-target scenarios while maintaining exceptional efficiency.
Jiayu Ding, Meilu Song, Xiaoyi Zhang et al.· 1 citation
This work proposes GroupForward, an instance-grouped feed-forward Gaussian splatting model that reconstructs geometry, appearance, instance structure, and semantics from sparse, unposed, and uncalibrated multi-view images and proposes a Referential Scene Reasoning Framework (RSRF) for complex 3D referring segmentation.
Qijian Tian, Zimeng Wu, Xuhong Wang et al.· 0 citations
Recent advances in 3D Gaussian Splatting (3DGS) have significantly advanced indoor scene reconstruction, a key capability for AR/VR, robotics, and autonomous systems. However, existing methods that rely on 2D monocular priors often produce artifacts such as holes and unnatural protrusions due to insufficient or weak constraints. To address this limitation, we propose a geometry-semantics co-regularization framework that jointly optimizes geometry and semantics within 3DGS. On the geometric side, we introduce a multi-level geometric optimization strategy with two key contributions. First, inspired by neural implicit surface representations, we design a dual implicit architecture that integrates Implicit Moving Least Squares (IMLS) and Signed Distance Functions (SDFs). This structure enables mutual reinforcement between the two representations via local geometric approximation and global normal constraints, improving fine-grained surface modeling. Second, we develop a view-dependent normal consistency constraint that combines 2D observations with SDF gradients to refine surface accuracy. On the semantic side, we develop a multi-view semantic consistency supervision to regularize the semantic distributions of Gaussian primitives, ensuring cross-view consistency for Gaussians corresponding to the same semantic category or instance. Extensive experiments on the Replica and MuSHRoom datasets demonstrate that our approach achieves state-of-the-art geometric reconstruction quality and rendering fidelity. Check out the project page https://sgo-gauss.github.io for the source code and more results.
Haihong Xiao, Jianan Zou, Yanan Zhang et al.· IEEE Transactions on Visuali...· 0 citations