Jul 2026· International Conference on Machine Vision and Applications· Vol 14270, pp. 142700A - 142700A-13· 0 citations· 40 references
Engineering
TL;DR
A 3D-GS-based framework that attaches a simplified Disney Bidirectional Reflectance Distribution Function (BRDF) and differentiable physically based rendering (PBR) shader to each surface Gaussian for interpretable material attributes delivers higher geometric accuracy, improved material realism, and real-time rendering performance suitable for applications such as Simultaneous Localization and Mapping (SLAM) and relighting.
Abstract
Efficient 3D reconstruction and high-fidelity novel-view synthesis under sparse-view, real-time constraints remain challenging. Implicit methods like Neural Radiance Fields (NeRF) depend on dense data and long optimization, while explicit 3D Gaussian Splatting (3D-GS) often loses geometric-photometric consistency and editable materials under sparse inputs. We propose a 3D-GS-based framework that attaches a simplified Disney Bidirectional Reflectance Distribution Function (BRDF) and differentiable physically based rendering (PBR) shader to each surface Gaussian for interpretable material attributes. A view-surface-aware selective densification strategy enhances sparse-region alignment, and a two stage training pipeline balances global consistency and surface fidelity, with 64 incident rays per Gaussian for accurate specular reflection. On MVImgNet and DTU dataset, our method achieves consistent improvements under 6-view settings: +0.4dB Peak Signal-to-Noise Ratio (PSNR), +0.5% Structural Similarity Index Measure (SSIM), and -5.6% Learned Perceptual Image Patch Similarity (LPIPS) on MVImgNet, and +0.6dB PSNR and +2.9% SSIM on DTU, compared with InstantSplat-XL. The proposed densification further improves PSNR by 0.7dB with only 20% higher training time. Overall, our framework delivers higher geometric accuracy, improved material realism, and real-time rendering performance suitable for applications such as Simultaneous Localization and Mapping (SLAM) and relighting.
Sparse view 3D reconstruction is commonly addressed with neural implicit surfaces or dense point-based representations such as Gaussian splatting. Surface-aware splatting methods improve extracted geometry through oriented primitives and regularization, while RadiosityGS incorporates differentiable light transport through a radiosity inspired finite-element surfel formulation. We propose a differentiable point rendering method based on opacity-bearing beta surfels. An opacity explicit adjoint light transport formulation provides gradients for surfel geometry and appearance parameters, allowing physically based light transport to constrain reconstruction. Across five synthetic objects reconstructed from ten posed views, our method achieves the lowest mean symmetric Chamfer distance among the evaluated baselines and reduces mean Chamfer distance by 28.5% relative to the strongest point-based baseline while using only 267 surfels on average, approximately ~161 fewer primitives. Directional Chamfer results further show improved accuracy and competitive completion relative to related point-based methods. These results show that, in the controlled direct illumination setting, compact beta surfels combined with transport-based optimization can recover surfaces without relying on the tens to hundreds of thousands of primitives used by the evaluated baselines.
M. K. Gjerde, J. B. Haurum, J. Frisvad et al.· 0 citations
3D reconstruction from sparse views is a challenging task in 3D computer vision. Recent studies on 3D Gaussian Splatting (3DGS) have achieved remarkable results with sparse views in novel view synthesis, yet reconstructing high-quality geometric surfaces from sparse views remains a challenge, due to the limited geometry clues and the discreteness of Gaussians. In this paper, we propose a novel 3DGS-based method for high-fidelity surface reconstruction from sparse views. Our key insight is to introduce a normal-guided depth propagation approach, which can extend depth information from high-confidence regions to constrain the depth in low-confidence areas. Additionally, we propose an abnormal depth edge-aware regularization to address depth discontinuities caused by the discreteness of Gaussians. Extensive experiments on DTU and Tanks-and-Temples datasets demonstrate that our method outperforms the state-of-the-art methods in sparse view surface reconstruction. Project page: https://hanl2010.github.io/DP-GS.
Liang Han, Bangcai Wei, Junsheng Zhou et al.· 0 citations
Reconstructing high-fidelity 3D scenes from sparse-views remains a central problem in generalizable neural rendering. Existing generalizable 3D Gaussian Splatting (3DGS) methods often exhibit geometric artifacts in sparse-view settings, since supervision based solely on 2D photometric losses cannot resolve depth and correspondence ambiguities. To address this issue, we propose MAC-Splat, a training framework built around direct 3D consistency supervision. MAC-Splat builds on the MASt3R geometric backbone and a frozen DINOv3 encoder to obtain semantically informed 2D correspondences, which serve as geometric anchors for 3D supervision. Using these anchors, we define the Multi-Attribute Consistency (MAC) loss. This objective jointly regularizes the 3D attributes of matched Gaussians, including their position, shape, and appearance, by enforcing agreement in a common world coordinate frame. The formulation is robust to outliers and respects the geometry of covariance matrices, which leads to stable training under sparse-view conditions. Experiments on ScanNet++ show that MAC-Splat outperforms strong baselines, with particularly large gains under different overlap regimes. In particular, it improves average PSNR over Splatt3R by more than 4.5 dB, reduces LPIPS, and maintains performance as the camera pose gap increases. These results indicate that a direct, multi-attribute 3D consistency objective, when combined with high-quality correspondences, is effective for addressing the ill-posed sparse-view reconstruction problem.
Jinqian Yang, Yichen Wu, Wanhua Li et al.· 1 citation
3D Gaussian Splatting (3DGS) has achieved remarkable success in novel view synthesis; however, reconstructions under sparse views often exhibit noticeable artifacts. While recent video diffusion models provide strong spatio-temporal priors for 3DGS restoration, directly fine-tuning them for restoration is suboptimal, as they lack awareness of the underlying multi-camera geometry, resulting in multi-view inconsistencies. In this work, we propose a novel 3D-aware video restoration framework designed to enhance the quality of sparse 3DGS reconstruction. Specifically, we construct a large-scale 3DGS video dataset to enable specialized fine-tuning. To bridge the gap between 2D video generation and 3D multi-view constraints, we introduce a camera-conditioned geometric prior. By using the first and last frames as boundary anchors and encoding the corresponding camera relationships, we explicitly inject spatial structure into the video generation pipeline. This boundary-anchored, camera-aware prior guides the network toward geometrically grounded restoration that remains coherent across viewpoints. Extensive experiments show that, among video-prior restoration methods, our approach attains the best pixel- and structure-level fidelity (PSNR/SSIM) and improves multi-view consistency, while remaining competitive in perceptual quality (LPIPS).
Xinhui Liu, Can Wang, Wei Jiang et al.· 0 citations
A semantic-guided 3D Gaussian splatting (3DGS) framework tailored to sparse-view industrial reconstruction was introduced, enabling robust reconstruction from limited viewpoints and offers a practical geometric foundation for automated inspection and remote equipment monitoring.
Boyang Li, Tianhan Gao, Zuan Gu et al.· Visual Computing for Industr...· 0 citations
The construction of high-fidelity twin geometric models is essential for advancing tunnel digital-twin technology. Image-based three-dimensional (3D) reconstruction, which directly infers 3D scene structures from visual semantics, has demonstrated considerable potential. However, most existing approaches rely on the conventional structure-from-motion (SfM) and multiview stereo (MVS) pipeline, which often suffers from point cloud voids and texture blurring in shield tunnel environments with low-texture segments and dim lighting. In addition, the inherent discreteness of point cloud representations makes subsequent denoising and optimization inefficient. To address these limitations, this study proposes a two-dimensional (2D) Gaussian splatting (2DGS) modeling approach integrated with a dynamic depth-aware masking mechanism. Image data are efficiently captured from a tunnel-longitudinal viewpoint, and sparse point clouds reconstructed via SfM are parameterized using 2DGS to enable adaptive geometric reconstruction under image supervision. Considering the linear and long-distance geometric characteristics of shield tunnels, a depth-aware dynamic masking strategy is introduced to guide the model to focus on near-field structural optimization under longitudinal viewing conditions. The optimized Gaussian model is rendered into depth maps at each camera viewpoint and fused using a truncated signed distance function to generate the final shield tunnel digital-twin geometric model. Experimental results from real tunnel engineering scenarios show that the proposed method achieves a geometric accuracy error of only 0.7% compared with SfM+MVS methods, significantly alleviates voids and local blurring artifacts, produces clearer segment contours, and reduces modeling time by approximately 37%. The proposed approach provides a novel technical paradigm for shield tunnel digital-twin geometric modeling.
Jinhua Qian, Weifeng Wei, Fei Xue et al.· Journal of computing in civi...· 0 citations