Skip to content
Book Open access

PixTex: Consistent 3D Texturing via Pixel-Space Multi-View Diffusion

Jul 2026 · International Conference on Computer Graphics and Interactive Techniques · pp. 1-12 · 1 citation · 76 references
Computer Science

TL;DR

PixTex is introduced, the first pixel-space multi-view diffusion framework for texture generation, which achieves substantially improved multi-view consistency, and proposes a novel consistency loss to explicitly guarantee multi-view coherence.

Abstract

Existing texture generation methods rely heavily on latent diffusion models, whose VAE-based spatial compression inherently limits fine-grained detail preservation and degrades pixel-level multi-view consistency. To address this limitation, we introduce PixTex, the first pixel-space multi-view diffusion framework for texture generation, which achieves substantially improved multi-view consistency. Operating directly in image space avoids latent compression, reduces inconsistencies introduced during latent-to-RGB upsampling, and preserves lossless pixel-level geometric guidance for accurate multi-view consistency. However, directly applying pixel-wise attention across multiple views is computationally prohibitive. To balance efficiency and fidelity, we adopt a coarse-to-fine consistency strategy: i) At a coarse patch level, we establish cross-view structural correspondence by employing 5D RoPE to correlate 2D patch coordinates with 3D world-space positions. ii) At the pixel level, a specialized 3D position-aware detailer further refines textural details based on patch features, ensuring fine-grained alignment unattainable by VAE-based methods. Additionally, we propose a novel consistency loss to explicitly guarantee multi-view coherence. Finally, we incorporate a pixel-space multi-view inpainting module to resolve self-occlusions and improve texture completeness. Extensive experiments demonstrate that our framework achieves state-of-the-art multi-view consistency, producing high-fidelity and seamless textures.

Read PDF

Similar papers

#diffusion models Preprint Sep 2026

SceneHI: High-Resolution 3D-Consistent Scene Texturing with Controllable Illumination

This work introduces an exact analytical pixel-to-texel mapping that aligns diffusion trajectories across multiple viewpoints, and utilizes High-Resolution Latent Textures as a persistent canvas for gradually denoised textures, while camera views perform the denoising steps in latent pixel space.

Athanasios Tragakis, M. Aversa, D. Ivanova et al. · 0 citations
Preprint Sep 2026

UltraTex: Unleashing 2K Multi-View Diffusion for 3D Texturing

This paper presents UltraTex, an efficient end-to-end framework for high-resolution multi-view diffusion-based 3D texturing and introduces Background Token Dropping, which removes background tokens before the DiT backbone, and Block-Sparse Attention, which reduces attention computation over the retained foreground sequ...

Yi-Bo Zhang, Ze Yuan, Nan Cao et al. · 0 citations
Preprint Sep 2026

Texture Space Material Diffusion

We present a method for generating high quality materials for 3D objects entirely in texture space. We finetune a video diffusion transformer for text-guided material generation, multi-view material generation, and material upscaling. Our key insight is to use the known projection from image space to texture space, ena...

Jacob Munkberg, Peter Kocsis, J. Hasselgren · 0 citations
Preprint Sep 2026

DirectUV: Image-Conditioned UV Texture Generation with Surface-Aware Positional Encoding

Generating high-quality UV textures for 3D meshes remains challenging. Multi-view projection pipelines suffer from occlusion and view inconsistency, and recent methods that generate textures directly in UV space still rely on auxiliary modules to supply 3D information, leaving the attention mechanism tied to UV-grid po...

Jian-Tao Lin, Ying-Jie Xu, Ming Sheng et al. · 0 citations
Open access Sep 2026

IP-ConTex: detail-consistent texture generation with image prompt

Textures are critical for enhancing the visual fidelity and diversity of three-dimensional (3D) models. Recently, generative models have significantly advanced texture generation. However, fine-grained control of the generation process remains challenging. Hence, we propose IP-ConTex, which is a novel image-guided text...

Lei Wang, Jie-Qing Feng · 0 citations
Preprint Aug 2026

GeoCache: Training-Free Acceleration of Multi-View Texture Diffusion via Geometric Delta Transport

Geometry-conditioned multi-view diffusion enables high-quality 3D texture generation, but its repeated per-view denoiser evaluations introduce substantial computational cost. Existing training-free accelerators primarily exploit temporal redundancy by reusing computation across denoising steps. In multi-view texturing,...

Haotang Li, Zhen-Yu Qi, Shaohan Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.