Skip to content

SARG-GS: semantic-augmented and residual-guided 3D Gaussian Splatting for sparse view synthesis

Jul 2026 · The Visual Computer · Vol 42 · 0 citations · 44 references

TL;DR

SARG-GS is proposed, a geometry-driven 3DGS framework tailored for sparse-view scenarios, comprising a Semantic Augmented Epipolar Fusion (SAEF) module and a Residual Guided Reprojection Compensation (RRC) module, which achieves superior structural completeness and rendering fidelity with as few as three input views.

View source

Similar papers

Open access Aug 2026

Semantic-guided 3D Gaussian splatting for sparse-view reconstruction in industrial digital twins

A semantic-guided 3D Gaussian splatting (3DGS) framework tailored to sparse-view industrial reconstruction was introduced, enabling robust reconstruction from limited viewpoints and offers a practical geometric foundation for automated inspection and remote equipment monitoring.

Boyang Li, Tianhan Gao, Zuan Gu et al. · 0 citations
Open access Jul 2026

VISTA-GS: MVS-Guided Virtual View Augmentation for Sparse-View 3D Gaussian Splatting

Abstract. 3D Gaussian Splatting (3DGS) has emerged as a leading technique for novel view synthesis (NVS), yet its performance degrades drastically under sparse-view conditions. While existing methods have sought to address this by incorporating accurate 3D geometry via Multi-View Stereo (MVS) or LiDAR priors, the view-dependent appearance parameters (i.e., spherical harmonics) remain exclusively optimized on the limited training views, leading to severe appearance overfitting. This is the fundamental reason why these geometry-enhanced methods still fail to generalize to out-of-distribution (OOD) viewpoints with large baselines, such as lane-changing trajectories in autonomous driving. To address this limitation, we propose VISTA-GS (Virtual Image Synthesis and Training Augmentation), a framework that synergizes MVS-based dense initialization with a physically-grounded virtual view augmentation strategy. Specifically, we position virtual cameras at strategic offsets around the original viewpoints and render virtual training images with binary validity masks via alpha-blending. By computing photometric losses exclusively within valid mask regions, VISTA-GS injects explicit angular constraints into the optimization process, effectively regularizing view-dependent appearance without relying on any external generative model. Experiments on the LLFF benchmark and a real-world LiDAR-scanned dataset demonstrate that our method achieves state-of-the-art NVS quality under sparse-view settings, with particularly significant improvements on challenging OOD viewpoints.

Hongsheng Huang, Yaxin Li, Shengjun Tang et al. · 0 citations
Preprint Aug 2026

CoMVS-GS: Collaborative Multi-View Stereo and 3D Gaussian Splatting for Surface Reconstruction

3D Gaussian Splatting enables efficient novel view synthesis, but accurate mesh reconstruction remains difficult in weakly observed and occluded regions, where Gaussian primitives may grow into unstable or geometrically inconsistent structures. We propose CoMVS-GS, a general surface reconstruction framework that combines Multi-View Stereo with Gaussian splatting. CoMVS-GS initializes Gaussian primitives from dense multi-view stereo points with pre-flattened scales and normal-aligned orientations, providing stronger geometric priors than sparse structure-from-motion initialization and reducing ambiguity during early optimization. It further introduces PatchMatch-3DGS Mutual Supervision, where Gaussian-rendered depths and normals initialize PatchMatch refinement, and refined PatchMatch depths supervise Gaussian optimization to improve weakly constrained geometry. For surface extraction, CoMVS-GS replaces truncated signed distance field voxel fusion with a Delaunay graph-cut meshing pipeline, reducing sensitivity to voxel resolution while preserving visibility-consistent surface evidence. Experiments on DTU, GauU-Scene V2, and MatrixCity show that CoMVS-GS remains competitive on object-level reconstruction and improves geometric accuracy and mesh compactness in outdoor scenes while maintaining high rendering quality.

Shihan Chen, Junjing Zhang, Qingsong Yan et al. · 0 citations
Preprint Jul 2026

MAC-Splat: Multi-Attribute Consistency for High-Fidelity Sparse-View Reconstruction

Reconstructing high-fidelity 3D scenes from sparse-views remains a central problem in generalizable neural rendering. Existing generalizable 3D Gaussian Splatting (3DGS) methods often exhibit geometric artifacts in sparse-view settings, since supervision based solely on 2D photometric losses cannot resolve depth and correspondence ambiguities. To address this issue, we propose MAC-Splat, a training framework built around direct 3D consistency supervision. MAC-Splat builds on the MASt3R geometric backbone and a frozen DINOv3 encoder to obtain semantically informed 2D correspondences, which serve as geometric anchors for 3D supervision. Using these anchors, we define the Multi-Attribute Consistency (MAC) loss. This objective jointly regularizes the 3D attributes of matched Gaussians, including their position, shape, and appearance, by enforcing agreement in a common world coordinate frame. The formulation is robust to outliers and respects the geometry of covariance matrices, which leads to stable training under sparse-view conditions. Experiments on ScanNet++ show that MAC-Splat outperforms strong baselines, with particularly large gains under different overlap regimes. In particular, it improves average PSNR over Splatt3R by more than 4.5 dB, reduces LPIPS, and maintains performance as the camera pose gap increases. These results indicate that a direct, multi-attribute 3D consistency objective, when combined with high-quality correspondences, is effective for addressing the ill-posed sparse-view reconstruction problem.

Jinqian Yang, Yichen Wu, Wanhua Li et al. · 1 citation
Jul 2026

Enhance 3D Gaussian splatting for dynamic scenes: integrating semantic and geometric consistency

3D Gaussian splatting (3DGS) provides an efficient and expressive scene representation by jointly modeling spatial geometry and appearance, which has led to significant advances in high-fidelity scene reconstruction and novel view synthesis. However, in real-world environments, occlusions from pedestrians or equipment often introduce blurriness, artifacts, and geometric distortions. To address these challenges, this paper proposes a robust 3DGS modeling method enhanced by semantic and geometric consistency. First, the self-supervised foundation model DINOv2 is utilized to extract high-dimensional semantic features, leveraging its superior generalization capabilities to assist in identifying potential dynamic regions. Second, monocular depth estimation and a depth residual mechanism are introduced to construct geometric consistency constraints, enabling the precise localization of areas that violate static assumptions. Finally, a progressive guided probability masking mechanism is designed; it employs an adaptive sigmoid function to achieve a ‘coarse-to-fine’ soft-constraint optimization, effectively mitigating the training instability inherent in traditional binary hard masks. Experimental results on the neural radiance fields (NeRF)-on-the-go, RobustNeRF, and self-collected datasets demonstrate that the proposed method effectively suppresses dynamic artifacts and improves reconstruction quality. The proposed approach achieves competitive or superior performance compared with 3DGS, SpotLessSplats, T-3DGS, and RobustSplat on standard image-quality metrics, including peak signal-to-noise ratio, structural similarity index measure, and learned perceptual image patch similarity.

Wen Zheng, Guo Bao, Wenda Wang et al. · 0 citations
Preprint Aug 2026

D^2-4DGS: Dual-Depth Guided Sparse-Camera 4D Gaussian Splatting

Dynamic 4D Gaussian Splatting has emerged as an efficient representation for dynamic novel view synthesis through explicit scene modeling and real-time rendering. However, existing methods typically require dense multi-view videos for sufficient geometric constraints, making capture expensive and limiting sparse-camera deployment. Reducing input views lowers acquisition cost but weakens geometry supervision, often causing missing structures and floating Gaussians. Depth priors provide geometric cues, yet no single source offers both dense coverage and reliable geometry. Monocular depth provides dense structure but is scale-ambiguous and locally biased, whereas multi-view geometric depth provides incomplete anchors consistent with the reconstruction coordinate system. To exploit their complementarity, we propose D$^2$-4DGS, a sparse-camera dynamic 4D Gaussian Splatting framework guided by dual-source depth priors. We align monocular estimates with valid multi-view geometric depths and verify their consistency to identify reliable geometric anchors. These verified anchors support consistency-aware pruning and depth supervision, while verified geometric depths and aligned mono-only estimates provide candidate geometry for densification in under-reconstructed regions. Finally, RGB-D joint optimization improves appearance fidelity and geometric consistency under sparse-view supervision. Across all nine dataset--view settings, D$^2$-4DGS achieves the highest PSNR, improving by 1.33 dB on average over the best competing method in each setting.

Jijian Zhao · 0 citations