Skip to content
Preprint

WilLaGS: Latent-Conditional 3D Appearance Fields for Robust Gaussian Splatting In-the-Wild

Aug 2026 · 0 citations · 53 references
Computer Science

TL;DR

A generative appearance model where a $\beta$-VAE learns a structured and continuous manifold of global appearance is introduced, Conditioned on the latent code, a 3D neural appearance field is constructed that generates dynamic Tri-Plane features to encode spatially-varying local illumination effects.

Abstract

3D Gaussian Splatting (3DGS) delivers real-time and high-fidelity rendering but remains challenged by unconstrained in-the-wild scenes, where drastic appearance variations and transient objects violate multi-view consistency. Existing methods are fundamentally limited by independent and discrete embeddings that struggle to capture continuous environmental changes or model spatially-varying local illumination. To address these limitations, we propose \textbf{WilLaGS}, a unified framework for robust 3D scene reconstruction and generative appearance synthesis under unconstrained settings. Specifically, we introduce a generative appearance model where a $\beta$-VAE learns a structured and continuous manifold of global appearance. Conditioned on the latent code, we construct a 3D neural appearance field that generates dynamic Tri-Plane features to encode spatially-varying local illumination effects. Furthermore, to suppress transient artifacts, we present a self-supervised perceptual masking mechanism that leverages a Teacher-Student (EMA) architecture to derive a stable scene consensus, robustly identifying inconsistent regions via perceptual discrepancies. Extensive experiments on multiple datasets demonstrate that \textbf{WilLaGS} achieves state-of-the-art performance in reconstruction quality and novel view appearance synthesis, while maintaining real-time rendering efficiency.

View source

Similar papers

Preprint Aug 2026

PoseAdapter: Dual-Stream 2.5D Controllable Image Generation for Complex Multi-Object Scenes

PoseAdapter, a lightweight framework for high-fidelity 2.5D controllable image generation, and a Context-Aware Dual-Stream Representation, to resolve the generative trade-off between strict instance isolation and global coherence.

Yu-Feng Chi, Hui-Min Ma, Fan Gao et al. · 0 citations
#diffusion models Preprint Sep 2026

Rethinking 3D Noise: Learning 3D-Aware Video Priors via Optimization-Free Morphological Perturbations

3D scene representations like NeRF and 3D Gaussian Splatting (3DGS) suffer severe artifacts in sparse-view settings. Recent generative 3D artifact fixers attempt to address this, but rely on paired corrupted and clean renders requiring costly, per-scene reconstructions across varying view configurations. While 2D image...

Onat Şahin, Mohammad Altillawi, George Eskandar et al. · 0 citations
Preprint Sep 2026

GAPS: Generative Active Pseudo-view Selection for Sparse-View 3D Gaussian Splatting

An alternating optimization framework that uses a pre-trained image diffusion model to generate geometrically consistent pseudo-views for additional 3DGS supervision and introduces Generative Active Pseudo-view Selection (GAPS) to balance reconstruction informativeness and generative reliability when choosing target vi...

Hong-Fei Zhu, Hao-Chen Deng, Si-Tao Zhang et al. · 0 citations
Preprint Aug 2026

Gaussian-JEPA: Joint-Embedding Predictive Learning for 3D Gaussian Splats

3D Gaussian Splatting (3DGS) represents 3D content with anisotropic primitives that jointly encode geometry and appearance. Fixed-budget encoders consume sampled observations of Gaussian assets, so the same object may be observed through different primitive realizations. Existing self-supervised methods mainly reconstr...

Bin Ren, Qi Ma, Yue Li et al. · 1 citation
#diffusion models Preprint Sep 2026

SceneHI: High-Resolution 3D-Consistent Scene Texturing with Controllable Illumination

This work introduces an exact analytical pixel-to-texel mapping that aligns diffusion trajectories across multiple viewpoints, and utilizes High-Resolution Latent Textures as a persistent canvas for gradually denoised textures, while camera views perform the denoising steps in latent pixel space.

Athanasios Tragakis, M. Aversa, D. Ivanova et al. · 0 citations
Preprint Aug 2026

FixAnything: 3D-Consistent Rendering Refinement via Video Generative Priors

This work presents FixAnything, a single model for fixing a wide range of rendering artifacts by repurposing a pretrained video generative model, leveraging its implicit multi-view priors with only minimal modification and lightweight finetuning.

Khiem Vuong, D. Ramanan, Srinivasa G. Narasimhan · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.