Skip to content
Preprint

GS-Voxel: Fitting-Free Structured Latents for Large-Scale 3DGS Generation

Aug 2026 · 1 citation · 45 references
Computer Science

TL;DR

The results show that GS-Voxel provides structured latents for pre-optimized aerial 3DGS reconstructions, with latent capacity that grows with the number of occupied voxels, with latent capacity that grows with the number of occupied voxels.

Abstract

Many scalable latent 3D generators operate on structured tensors, whereas pre-optimized 3D Gaussian Splatting (3DGS) reconstructions are unordered, spatially irregular, and vary widely in primitive count. We present GS-Voxel, a fitting-free structured latent framework, and evaluate it for large-scale aerial 3D Gaussian scene generation. GS-Voxel deterministically converts a compatible pre-optimized 3DGS reconstruction into sparse active voxels without additional per-scene optimization, retaining the sub-voxel positions and rendering attributes of the selected primitives. A GS-specific factorized VAE then separately encodes voxel geometry and local Gaussian attributes into sparse 3D latents whose size grows with the number of occupied voxels rather than being limited by a fixed scene-wide primitive count. We train image-conditioned flow models in the GS-Voxel latent space to generate aerial 3DGS scenes. A key application enabled by GS-Voxel is large-area scene generation: overlap-aware tiled inference extends synthesis beyond a single training crop conditioned on satellite-view images. Our results show that GS-Voxel provides structured latents for pre-optimized aerial 3DGS reconstructions, with latent capacity that grows with the number of occupied voxels.

View source

Similar papers

Preprint Sep 2026

FILIGREE3D: Scaling Sparse Latent Flow Matching for Ultra-High-Resolution Image-to-3D Generation

Scaling image-to-3D generation to ultra-high resolutions requires controlling rapidly growing computational costs without sacrificing fine geometric detail. We present \textbf{Filigree3D}, a sparse latent flow-matching framework that generates 3D geometry from a single image at voxel resolutions up to $2048^3$, with st...

Hong-Ji Li, Xin-Ran Yang, Xiu-Chao Wu et al. · 0 citations
Preprint Aug 2026

VoxStruct3D: Structure-Leading Flow Matching for Voxel-Space 3D MRI Synthesis

High-fidelity 3D MRI synthesis requires both globally coherent anatomy and fine-grained voxel-level detail. Although latent diffusion makes volumetric generation tractable, its image autoencoder introduces a reconstruction bottleneck that can limit the fine detail recoverable in the final volume. We present VoxStruct3D...

Fang Li, Yang Gao, Shihao Zou et al. · 0 citations
#machine learning Preprint Sep 2026

Sparse auto-regressive modeling for scene generation from multi-view images

This work introduces SPAR3S, a sparse voxel-aligned 3D latent generative model for conditional scene completion without requiring ground-truth 3D data for supervision, and trains a masked autoregressive transformer that jointly models voxel occupancy and latent token values, enabling efficient and spatially consistent...

Thomas Lucas, Maxime Pietrantoni, Philippe Weinzaepfel et al. · 1 citation
Preprint Aug 2026

Visibility-Guided Structured Measure Flow for Class-Conditioned 3D Gaussian Generation

3D Gaussian Splatting (3DGS) has made real-time, high-fidelity 3D rendering practical, yet turning this explicit representation into a native generative space remains an open challenge. Directly generating 3DGS objects is difficult because Gaussian primitives are unordered, variable-sized, locally dense, and highly sen...

Yi-Zhao Wang · 0 citations
Preprint Aug 2026

Fusion-Aware Direct 3D Gaussian Generation with Structured Patch Latent Flows

A structure-aware rectified flow model with patch-position conditioning, global-local coupled velocity prediction, and density-aware velocity weighting is designed, enabling direct latent generation of class-conditioned 3DGS objects within seconds.

Yi-Zhao Wang, Jing-Bo Wang, Guan-Tao Zhang · 0 citations
Preprint Sep 2026

PLSR: Progressive and Localized Super-Resolution of 3D Objects via Localized Latent Voxel Diffusion

High-resolution 3D asset generation is vital in various 3D applications. Existing state-of-the-art diffusion-based models remain constrained by fixed resolutions, limiting their ability to produce details. In this paper, we tackle the challenge of generating more detailed, higher-resolution 3D objects by introducing a...

Yu-Xin Liu, Minshan Xie, Jia-Wen Liang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.