Skip to content
Preprint

UltraTex: Unleashing 2K Multi-View Diffusion for 3D Texturing

Sep 2026 · 0 citations · 39 references
Computer Science

TL;DR

This paper presents UltraTex, an efficient end-to-end framework for high-resolution multi-view diffusion-based 3D texturing and introduces Background Token Dropping, which removes background tokens before the DiT backbone, and Block-Sparse Attention, which reduces attention computation over the retained foreground sequence.

Abstract

High-quality texture generation is essential for creating realistic and production-ready 3D assets. Recent multi-view diffusion methods have shown promising results for image-guided 3D texturing, but they are typically constrained to low operating resolutions such as 512 or 768, making it difficult to preserve high-frequency details from high-resolution reference images. Scaling this paradigm to 2048 resolution is computationally prohibitive, as the unified multi-view sequence exceeds 212K tokens and incurs excessive memory and latency. In this paper, we present UltraTex, an efficient end-to-end framework for high-resolution multi-view diffusion-based 3D texturing. Our key observation is that object-centric multi-view renderings contain two major sources of redundancy: background-induced sequence redundancy and sparse token interactions within the foreground. To address them, we introduce Background Token Dropping, which removes background tokens before the DiT backbone, and Block-Sparse Attention, which reduces attention computation over the retained foreground sequence. To enable efficient foreground-only inference while avoiding reconstruction artifacts, we further design Foreground-Aware VAE Decoding to ensure the quality of the final high-resolution views. To satisfy the demanding data requirements of 2K-resolution multi-view diffusion training, we construct G-buffer TexVerse, a large-scale, ultra-high-resolution multi-view rendering dataset covering over 268,000 3D assets. Extensive experiments show that UltraTex generates visually faithful textures with rich fine-grained details, while substantially improving efficiency, achieving $20.6\times$--$91.1\times$ training speedup and $22.3\times$--$74.6\times$ end-to-end inference speedup over the baseline on common samples in our dataset. Code and data is at https://yiboz2001.github.io/UltraTex.

View source

Similar papers

#diffusion models Preprint Sep 2026

SceneHI: High-Resolution 3D-Consistent Scene Texturing with Controllable Illumination

This work introduces an exact analytical pixel-to-texel mapping that aligns diffusion trajectories across multiple viewpoints, and utilizes High-Resolution Latent Textures as a persistent canvas for gradually denoised textures, while camera views perform the denoising steps in latent pixel space.

Athanasios Tragakis, M. Aversa, D. Ivanova et al. · 0 citations
Preprint Aug 2026

EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing

This work proposes EditBridge, a diffusion bridge framework for efficient ultra high-resolution editing that achieves high-fidelity editing with superior perceptual quality at resolutions up to 4K, delivering 3.6--8.4$\times$ speedup at 2K and enabling practical 4K editing in 61 seconds.

Jiayi Song, Shijie Huang, Fang-Tai Wu et al. · 0 citations
Preprint Sep 2026

Deformable 2D Gaussian Splatting for Efficient 4K Video Compression

This work proposes a real-time video compression framework that represents and compresses a Group of Pictures using a coarse-to-fine multi-scale 2D Gaussian Splatting structure coupled with a lightweight deformation network and demonstrates the potential of Gaussian Splatting as a practical solution for efficient high-...

Chen-Hao Zhang, Feng-Qing Zhu · 0 citations
Preprint Sep 2026

LiteTex-GS: Fast and Lightweight Texturing for Gaussian Splatting

Gaussian Splatting has enabled real-time novel view synthesis, but its tightly coupled geometry and appearance representation often require a large number of primitives to reproduce high-frequency texture details, leading to substantial memory and optimization costs. Recent textured 2D Gaussian methods alleviate this l...

Zhi-Wei Li, Yi-Jia Guo, Yi-Shi Lu et al. · 0 citations
Conference Aug 2026

ENEA-GS: Enhancing Single-Image 3D Gaussian Splatting with Normal-guided Texture Smoothness and Entropy-based Alpha Regularization

Recent advances in 3D content generation have demonstrated the effectiveness of optimization-based frameworks such as DreamGaussian, which combine 3D Gaussian Splatting (3DGS) with diffusion-guided Score Distillation Sampling (SDS) to efficiently synthesize 3D assets from a single image. Compared with earlier NeRF-base...

Nam-Quan Nguyen, Lam-Huy Nguyen, Minh-Triet Tran · 0 citations
Preprint Aug 2026

PoseAdapter: Dual-Stream 2.5D Controllable Image Generation for Complex Multi-Object Scenes

PoseAdapter, a lightweight framework for high-fidelity 2.5D controllable image generation, and a Context-Aware Dual-Stream Representation, to resolve the generative trade-off between strict instance isolation and global coherence.

Yu-Feng Chi, Hui-Min Ma, Fan Gao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.