Jul 2026· International Conference on Computer Graphics and Interactive Techniques· pp. 1-12· 1 citation· 76 references
Computer Science
TL;DR
PixTex is introduced, the first pixel-space multi-view diffusion framework for texture generation, which achieves substantially improved multi-view consistency, and proposes a novel consistency loss to explicitly guarantee multi-view coherence.
Abstract
Existing texture generation methods rely heavily on latent diffusion models, whose VAE-based spatial compression inherently limits fine-grained detail preservation and degrades pixel-level multi-view consistency. To address this limitation, we introduce PixTex, the first pixel-space multi-view diffusion framework for texture generation, which achieves substantially improved multi-view consistency. Operating directly in image space avoids latent compression, reduces inconsistencies introduced during latent-to-RGB upsampling, and preserves lossless pixel-level geometric guidance for accurate multi-view consistency. However, directly applying pixel-wise attention across multiple views is computationally prohibitive. To balance efficiency and fidelity, we adopt a coarse-to-fine consistency strategy: i) At a coarse patch level, we establish cross-view structural correspondence by employing 5D RoPE to correlate 2D patch coordinates with 3D world-space positions. ii) At the pixel level, a specialized 3D position-aware detailer further refines textural details based on patch features, ensuring fine-grained alignment unattainable by VAE-based methods. Additionally, we propose a novel consistency loss to explicitly guarantee multi-view coherence. Finally, we incorporate a pixel-space multi-view inpainting module to resolve self-occlusions and improve texture completeness. Extensive experiments demonstrate that our framework achieves state-of-the-art multi-view consistency, producing high-fidelity and seamless textures.
This work introduces an exact analytical pixel-to-texel mapping that aligns diffusion trajectories across multiple viewpoints, and utilizes High-Resolution Latent Textures as a persistent canvas for gradually denoised textures, while camera views perform the denoising steps in latent pixel space.
Athanasios Tragakis, M. Aversa, D. Ivanova et al.· 0 citations
This paper presents UltraTex, an efficient end-to-end framework for high-resolution multi-view diffusion-based 3D texturing and introduces Background Token Dropping, which removes background tokens before the DiT backbone, and Block-Sparse Attention, which reduces attention computation over the retained foreground sequ...
We present a method for generating high quality materials for 3D objects entirely in texture space. We finetune a video diffusion transformer for text-guided material generation, multi-view material generation, and material upscaling. Our key insight is to use the known projection from image space to texture space, ena...
Jacob Munkberg, Peter Kocsis, J. Hasselgren· 0 citations
Generating high-quality UV textures for 3D meshes remains challenging. Multi-view projection pipelines suffer from occlusion and view inconsistency, and recent methods that generate textures directly in UV space still rely on auxiliary modules to supply 3D information, leaving the attention mechanism tied to UV-grid po...
Jian-Tao Lin, Ying-Jie Xu, Ming Sheng et al.· 0 citations
Textures are critical for enhancing the visual fidelity and diversity of three-dimensional (3D) models. Recently, generative models have significantly advanced texture generation. However, fine-grained control of the generation process remains challenging. Hence, we propose IP-ConTex, which is a novel image-guided text...
Lei Wang, Jie-Qing Feng· Visual Computing for Industr...· 0 citations