Skip to content

4DHumanDiff: Direct Text-to-4DGS Generation for Consistent 360-Degree Dynamic Humans

Jul 2026 · arXiv.org · Vol abs/2607.27634 · 0 citations · 127 references
Computer Science

TL;DR

4DHumanDiff is presented, a diffusion framework that directly generates dynamic humans represented by 4D Gaussian Splatting from text prompts and achieves better temporal and multi-view consistency, and reduces inference time by more than 10x.

Abstract

Generating high-quality 360-degree dynamic human assets from text prompts is challenging. Existing methods usually synthesize monocular or multi-view videos first and then fit a 4D representation, which is expensive and often causes incomplete geometry or view-inconsistent renderings. We present 4DHumanDiff, a diffusion framework that directly generates dynamic humans represented by 4D Gaussian Splatting (4DGS) from text prompts. By modeling the structured 4D representation space end-to-end, 4DHumanDiff avoids video pre-generation and per-scene reconstruction, making it better suited for view-consistent and temporally coherent asset generation. The model uses a 3D U-Net backbone with temporal attention for motion-aware generation. We further construct a large-scale text-to-4DGS dataset with 60,000 high-quality pairs, and introduce 2D regularization and training-free 4D interpolation to improve rendering quality and motion smoothness. Experiments show that 4DHumanDiff generates consistent 360-degree dynamic humans within one minute, achieves better temporal and multi-view consistency, and reduces inference time by more than 10x.

View source

Similar papers

Open access Aug 2026

TGRHuman: Text-Guided Realistic 3D Human Generation via Diffusion Renderer

TGRHuman, a novel approach for generating realistic 3D humans from text that decouples geometry and texture generation to alleviate the issues commonly encountered in NeRF-based methods, outperforms existing text-to-3D human methods in both geometry and texture quality.

Muxin Zhang, Chao-Hui Yu, Yuanwang Yang et al. · 0 citations
Preprint Aug 2026

4DAnyone: Create Anyone in 4D from a Casual Monocular Video

4DAnyone outperforms prior methods in both novel-view video quality and downstream 4DGS reconstruction, with robust in-the-wild generalization.

Yudong Jin, Tao Xie, Qihang Zhang et al. · 0 citations
Preprint Aug 2026

Block3D: Efficient Text-to-3D Generation via Block-Wise Diffusion

Block3D is proposed, a block-wise diffusion framework that partitions the discrete shape-token sequence into contiguous blocks, generates the blocks autoregressively, and jointly denoises all tokens within the current block and introduces confidence-guided intra-block correction, which revises low-confidence tokens bef...

Bowen Cui, Weijie Wang, Zeyu Zhang et al. · 1 citation
Preprint Aug 2026

DiGS-Avatar: Single-Image Animatable 3D Human Reconstruction via UV-Space Diffusion

DiGS-Avatar is proposed, which reformulates this task as an efficient, diffusion-based UV-latent completion task, ensuring 3D consistency by design, and introduces a teacher-student framework where a multi-view teacher provides geometrically aligned pseudo-ground-truth latents to supervise a single-view diffusion stude...

Jiakun Li, Li Fang, Hao Zhu et al. · 0 citations
Preprint Sep 2026

UltraTex: Unleashing 2K Multi-View Diffusion for 3D Texturing

This paper presents UltraTex, an efficient end-to-end framework for high-resolution multi-view diffusion-based 3D texturing and introduces Background Token Dropping, which removes background tokens before the DiT backbone, and Block-Sparse Attention, which reduces attention computation over the retained foreground sequ...

Yi-Bo Zhang, Ze Yuan, Nan Cao et al. · 0 citations
Conference Aug 2026

4D-GAIA:Depth-Guided 4D Gaussian Representation for Dynamic Visual Data Analytics

Dynamic visual data, such as monocular videos of moving humans and changing environments, contain rich spatio-temporal information but are often large, redundant, and difficult to represent in a compact and structured form. 4D Gaussian Splatting has recently emerged as an explicit representation for dynamic scene recon...

Junita Sari, Nicolas Slat, Di Jiang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.